The GPU OS is the firmware in CPU2's 16 kB ROM. It boots the video subsystem, then
runs one interrupt-driven render pass per VSYNC, executing the instruction list CPU1
left in ping-pong RAM. Source: roms/gpu_os.s.
Status: pre-1.0. The ROM reports
V0.9on its diagnostic screen. Several features are still outstanding and nothing has run on real hardware yet — only inmadsimand the Verilator sim. Call this v0.8→v0.9: the instruction set and the PPRAM protocol are stable enough to write cartridges against, but they are not frozen until a board boots.
Comb rendering and vertical (TATE) text below are both implemented and shipped. What is
still designed but not built — DOT_CIRCLE16 and the TATE rotate-90 asset script —
is in MAD65_future_features.md.
$0000–$77FF RAM scratch, stack, tables, graphics pool
$7800–$7FFF PPRAM ping-pong RAM 2 kB — inter-CPU communication slot
$8000–$BFFF VRAM-image active write buffer, hardware-switched each VSYNC
$C000–$FFFF ROM (read) / VRAM-background (write)
VRAM note: only one VRAM buffer is accessible to the GPU at any time. Hardware switches buffers each VSYNC automatically. The GPU never knows which physical buffer it currently holds — this is the root of the two-frame background rule below.
$0000–$77FF, 30 kB)#$0000–$00FF Zero page 256 B see below
$0100–$01FF Stack 256 B 65C02 hardware stack
$0200–$02FF Scratch 256 B per-opcode working buffers
$0300–$06FF Sprite definition table 1024 B 4 pages × 256 entries
$0700–$0EFF Tile bank 2048 B 256 user tiles × 8 B
$0F00–$0FFF (reserved) 256 B
$1000–$77FF Graphics pool 26624 B 26 kB — sprite/tile bitmaps, uploaded via LOAD
The $0200 scratch page is shared and overlapping by design — only one opcode
handler runs at a time, so the sprite row spans (SPR_BMP $0200, SPR_OVL $0210),
draw_text_core's shift buffer and op_tile's glyph-pointer tables all reuse it.
$0000–$00FF)#$00 GPU_STATUS ZP mirror of the PPRAM status byte
$01 NO_CPU_COUNT consecutive frames without CPU_READY (diag mode trips at $40)
$02 VIDEO_REG_SHADOW software shadow of the write-only VIDEO_REG
$03–$04 PPWP instruction-list cursor (PPRAM, or DIAG_LIST in diag mode)
$05–$06 PTR universal 16-bit VRAM / RAM pointer
$07–$0A PX_X, PX_Y full-res pixel coordinates, 16-bit each
$0B–$0C PX_TMP address-arithmetic scratch
$0D–$0E GLYPH_PTR current glyph row in the font
$0F DRAW_BASE_HI VRAM target high byte: $80 = image, $C0 = background
$10–$14 TEXT_* column, row, scroll, visible length, glyph-row counter
$15–$21 LN_* line state: X1,Y1,X2,Y2, DX,DY, SX,SY, ERR, CNT, MASK,
ROWSTEP (signed 16-bit, shared by line_core + dotline_core)
$22–$33 SPR_* sprite blitter: width, rows, overlay flag, bit offset, shift
count + direction, span, column, clip flag, row counter,
source ptr, VRAM ptr, id, top-skip, stride, composite scratch
$34–$49 CIR_* DOT_CIRCLE: centre, radius, midpoint walker, decision
variable, per-group column/row values, masks and validity
$4A BOOT_BLINDED one-shot — boot's BLINDER is still up; drop it on frame 1
$4B–$4E SPR_COMB, SPR_D, SPR_VSTEP comb-sprite ($51) state
$4F–$52 SPR_CT0/CT1 comb-sprite 16-bit scratch
$53–$64 L16_* LINE16 / chain state: 16-bit endpoints, deltas, error,
step directions, the chain flag and the segment counter
$65 FRAME_BUSY set while a command list is being dispatched — the
frame-overrun detector (see *Per-frame ISR*)
$66 OVERRUN_CNT overrun frames since boot, saturating at 255
$67–$A7 PG_* polygon ($4C–$4E): centre, angle, scale, the rotate-and-scale
matrix and its signs, the vertex cursor and outcodes, the
bisection clipper's working points, and the multiply operands
$A8–$C3 CI_* CIRCLE16 ($4F): signed-16 full-res centre, radius, midpoint
walker and decision variable, the clip flag, and the per-group
column/row values, masks, validities and row base
$C4–$FF free general ROM working variables
VIDEO_REG_SHADOWis the only record of the currentVIDEO_REGstate (the hardware register is write-only). All code must read the shadow and write to both — never writeVIDEO_REGwithout updating the shadow.
The first byte at $7800 is the status code byte (PP_STATUS); each side reads it to learn the
other's state. The GPU does not process PPRAM content until the CPU has set
CPU_READY. Bytes $7801 onward form the instruction list — a fixed opcode
followed by zero or more argument bytes, terminated by WAI ($00).
| Code | Symbol | Direction | Meaning |
|---|---|---|---|
$A0 |
CPU_BOOTING |
CPU → GPU | CPU is initialising |
$A1 |
CPU_WORKING |
CPU → GPU | CPU is processing a frame |
$A2 |
CPU_READY |
CPU → GPU | CPU has finished writing |
$B0 |
GPU_BOOTING |
GPU → CPU | GPU is initialising |
$B1 |
GPU_WORKING |
GPU → CPU | GPU is rendering |
$B2 |
GPU_READY |
GPU → CPU | GPU has finished rendering |
The CPU does not zero PPRAM per frame: it overwrites the list in place from
$7801 and re-terminates it with WAI; stale bytes past the WAI are never read.
| Opcode | Mnemonic | Arguments (bytes following opcode) | Description |
|---|---|---|---|
$00 |
WAI |
— | End of instruction list |
$08 |
GPU_LED |
pattern |
Write bit pattern to LED_GPU_REG |
$10 |
COPY_DIS_ON |
— | Set COPY_DIS in VIDEO_REG |
$11 |
COPY_DIS_OFF |
— | Clear COPY_DIS |
$12 |
BG_REG_ON |
— | Set BG_REG |
$13 |
BG_REG_OFF |
— | Clear BG_REG |
$14 |
BLINDER_ON |
— | Set BLINDER (blank screen) |
$15 |
BLINDER_OFF |
— | Clear BLINDER |
$20 |
CLEAR_BG |
— | Clear VRAM-background $C000–$FFFF. Two-frame rule applies |
$30 |
LOAD |
page_MSB, then 256 data bytes |
Copy 256 bytes into one page (page_MSB << 8). Valid: $02–$77 (RAM), $80–$BF (VRAM-image), $C0–$FF (VRAM-bg, two-frame rule). Blocked: $00, $01, $78 — data consumed but skipped |
$31 |
RECT_BG |
XB, Y_LSB, Y_MSB, WB, H, GAP, then WB*H data bytes |
Byte-aligned rectangle onto VRAM-background. Source row r lands on row Y + r·(GAP+1) — the comb step, GAP=0 = solid. XB/WB are BYTE columns (0–49 / 1–50). Rows past 299 are sunk, not clipped: their payload is still consumed. Two-frame rule applies |
$32 |
RECT_BG_RLE |
same header, then an RLE stream | The same, payload compressed. See The transport block below |
$33 |
LOAD_RLE |
page, LEN_LSB, LEN_MSB, then an RLE stream |
$30 with a compressed payload and a 16-bit length, so one command can span several destination pages. Same three blocked pages as $30, and a blocked page still consumes the stream |
$40 |
PIXEL |
X_LSB, X_MSB, Y_LSB, Y_MSB |
Pixel on VRAM-image, full-res |
$41 |
(removed) | — | PIXEL_BG — retired: a per-bit set needs read-modify-write, and VRAM-background is write-only (reads return ROM) |
$42 |
LINE |
X1, Y1, X2, Y2 |
Solid line on VRAM-image, half-res coords |
$43 |
LINE16 |
X1_LSB, X1_MSB, Y1_LSB, Y1_MSB, X2_LSB, X2_MSB, Y2_LSB, Y2_MSB |
Solid line on VRAM-image, full-res endpoints (0–399 / 0–299). Was LINE_BG through v0.8; the number was reused |
$44 |
DOT_LINE |
X1, Y1, X2, Y2 |
Dotted line, image only |
$45 |
DOT_LINES |
N, x0,y0, … , xN,yN |
Connected chain of N dotted segments (polyline), image only |
$46 |
DOT_PIXEL |
X, Y |
One pixel at half-res coords, image only |
$47 |
DOT_PIXELS |
N, x0,y0, … , xN-1,yN-1 |
N independent half-res pixels in one dispatch, image only |
$48 |
DOT_CIRCLE |
CX, CY, R |
Dotted outline circle, half-res, image only. Centre must be on-screen |
$4A |
LINES |
N, x0,y0, … , xN,yN |
Connected chain of N solid half-res segments — the solid twin of $45 |
$4B |
LINES16 |
N, x0.16,y0.16, … , xN.16,yN.16 |
Connected chain of N solid full-res segments |
$49 |
HDOT_LINE |
XB, Y, NB |
Byte-aligned horizontal dotted rule, image only. Self-clipping |
$4C |
DOT_POLYGON |
CX.16, CY.16, ANGLE, SCALE, N, dx0,dy0, … , dxK-1,dyK-1 |
Figure around a signed-16 centre, rotated by ANGLE and scaled by SCALE on the GPU, always clipped. N bit 7 = OPEN (polyline) / 0 = CLOSED. Dotted, half-res. 2K + 8 bytes, K = N & $7F |
$4D |
POLYGON |
same as $4C |
The solid half-res twin |
$4E |
POLYGON16 |
same as $4C |
The solid full-res twin: centre and offsets are read on the 400 × 300 grid |
$4F |
CIRCLE16 |
CX_LSB, CX_MSB, CY_LSB, CY_MSB, R |
Solid circle outline around a signed-16 full-res centre that may be off screen, always clipped. R in full-res pixels, 0–255. No ANGLE/SCALE — see below |
$50 |
SPRITE |
SPR_ID, X_LSB, X_MSB, Y_LSB, Y_MSB |
Sprite on VRAM-image, signed-16 top-left pixel, clipped on all four edges |
$51 |
COMB_SPRITE |
SPR_ID, X_LSB, X_MSB, Y_LSB, Y_MSB, GAP (0–15) |
Comb ("ghost") sprite — same sprite, same slot, but source row r lands on screen row Y + r·(GAP+1); the GAP rows between are never written, so the background shows through. Bitmap-only: the overlay plane is ignored. GAP=0 = a plain sprite without its overlay |
$60 |
TEXT |
X (0–49), Y (0–36), scroll (0–7), …string…, $00 |
Text on VRAM-image. Last character is never drawn — it is the scroll guard |
$61 |
TEXT_BG |
same as $60 |
Text on VRAM-background. Two-frame rule applies |
$62 |
VTEXT |
X (0–36), Y (0–49), scroll (0–7), …string…, $00 |
Vertical (TATE) text on VRAM-image — for a monitor stood on its side. Same (X, Y) order as TEXT; the 37 × 50 grid is its transpose. Every character is drawn, and one running past cell 36 is clipped per glyph row on the two guard cells |
$63 |
VTEXT_BG |
same as $62 |
Vertical text on VRAM-background. Two-frame rule applies |
$70 |
TILE |
X (0–49), Y (0–36), scroll (0–7), …tile ids…, $00 |
Tile string on VRAM-image |
$71 |
TILE_BG |
same as $70 |
Tile string on VRAM-background. Two-frame rule applies |
The CPU-side builders map 1:1 onto these opcodes — see
MAD65_CPU_OS.md.
Three line families, each offering the same four calls. Pick the family for the look and precision; the call within it for how coordinates behave at the screen edge.
| single, clamped | single, clipped | chain (polyline) | figure/polyline (rotated+scaled, GPU-side) | |
|---|---|---|---|---|
| dotted, half-res | gpu_dotline $FF1B → $44 |
gpu_dotline_clip $FF93 → $44 |
gpu_dotlines $FF1E → $45 |
gpu_dotpolygon $FFD2 → $4C |
| solid, half-res | gpu_line $FF15 → $42 |
gpu_line_clip $FFC3 → $42 |
gpu_lines $FFC6 → $4A |
gpu_polygon $FFD5 → $4D |
| solid, full-res | gpu_line16 $FFC0 → $43 |
gpu_line16_clip $FFBD → $43 |
gpu_lines16 $FFC9 → $4B |
gpu_polygon16 $FFD8 → $4E |
project_raw/mesh_xform_raw) and
cuts the segment against the screen rectangle, keeping its slope. ~630–855 cycles
trivially inside, ~6 500 when it must cut and divide, ~250 with nothing emitted when
wholly outside — earns its keep on off-screen geometry.$4B; ~245 for $4A), more valuable the shorter the edges. PPRAM drops from 9N
to 4N+6 bytes ($4B) or 5N to 2N+3 ($4A). Chains clamp only — a polyline that may
leave the screen must go edge-by-edge through *_clip, losing the joint saving.N's bit 7 picks CLOSED (a figure, the default) or OPEN (a
polyline through the same matrix). See The polygon family below.All line opcodes take half-resolution coordinates (0–199 × 0–149, one byte each),
doubled internally to the 400×300 framebuffer — halves PPRAM usage, keeps Bresenham in
8-bit arithmetic. DOT_LINE/DOT_LINES skip every other row so half-res stepping is
invisible there.
LINE ($42) draws at FULL resolution — only the endpoints are half-res (doubled
once); Bresenham then steps one full-res pixel at a time (8-connected, 1 px thick,
2·max(DX,DY)+1 pixels). LINE16 ($43) is the same algorithm with signed-16
endpoints instead of doubled half-res ones, so endpoints land on any pixel rather than
only even ones — the win is motion: with $42 a slowly rotating/drifting line snaps
in 2 px jumps (13 distinct lines out of a 25-frame 10° sweep vs. $43's 24). Cost: 8
PPRAM bytes instead of 4, and +5–8% cycles up to 255 px on the major axis (a 16-bit
Bresenham error term beyond that, +32–37% — op_line16 picks between two octant-loop
sets off the major delta's high byte). Use $42 for anything static, $43 when
endpoints move.
Feeding
$43/$4Bneeds the full-res projection — this is a real trap. The CPU vector library's ordinaryproject/mesh_xform(and*_raw) produce half-res coordinates; doubling one to get a full-res value can only ever yield an EVEN coordinate — exactly$42's quantisation, so a$43fed that way costs 8 extra PPRAM bytes to draw precisely what$42draws. Project throughmesh_xform_raw16($FFCC) orproject_raw16($FFCF) instead — same perspective, but the quotient lands on the 400 × 300 grid and comes out odd about half the time. Cost and caveats:MAD65_CPU_OS.md.
| Opcode | Rendering | Speed | Best for |
|---|---|---|---|
LINE ($42) |
Solid — full-res Bresenham, 8-connected, 1 px thick; H/V runs byte-batched | baseline; H/V lines ~6–8× faster via byte fill | UI frames, borders, axes, bars — anything that must look continuous |
LINE16 ($43) |
The same line, but the endpoints are full-res (16-bit) | +5–8 % over LINE up to 255 px on the major axis, +32–37 % past it |
Vectors whose ends move: rotating wireframes, needles, sweeps — anything that snaps visibly at 2 px. Project it with mesh_xform_raw16, or it is just $42 with more bytes |
DOT_LINE ($44) |
Dotted — one dot per half-res step, on even coords | ~1.8× faster than LINE |
One-off diagonal edges |
DOT_LINES ($45) |
Dotted chain, shares vertices | ~1.4–1.5× faster than the same edges sent separately, on top of DOT_LINE's gain |
Wireframes and any connected path |
DOT_CIRCLE ($48) |
Dotted outline, midpoint algorithm, 8-way symmetry | ~0.6 ms at R≈50, ~0.25 ms at R≈25 |
Reticles, radar rings, gauges, wireframe nodes |
HDOT_LINE ($49) |
Dotted horizontal rule — constant $AA byte fill |
~10–14× faster than the same span as a horizontal DOT_LINE (~13 cyc/byte vs ~194) |
HUD separators, grid/ruler lines, table rules, dotted underlines |
All of them are image-only — there is no _BG variant of any line or pixel op.
Setting an individual bit requires a read-modify-write and the VRAM-background window
is write-only. Background layers are built from whole-byte writes: LOAD, TEXT_BG,
TILE_BG, CLEAR_BG.
DOT_LINES shares each interior vertex between adjacent segments, so the expensive
start-address computation runs once per chain; branching models are sent as several
chains plus the occasional lone DOT_LINE.
DOT_CIRCLE restrictions#CX > 199 or CY > 149 → the whole circle is
skipped. An off-screen centre is a caller error, not a request to draw a far-off arc.
R = 0 also draws nothing.HDOT_LINE encoding and self-clipping#Three guarantees collapse the whole line to a constant memory fill: the row never
changes; every dot lands on an even full-res column so a full byte of dots is always
$AA; and both start and length are byte-aligned, so no mask is needed. The renderer
points at the start byte and stores $AA to NB consecutive VRAM bytes.
| Arg | Meaning | Range |
|---|---|---|
XB |
start byte column (full-res col = XB·8) |
0–49 |
Y |
half-res row — drawn on full-res row 2·Y |
0–149 |
NB |
run length in bytes (1 byte = 8 full-res columns) | 1–50 |
It is one of only two ops that validate their own arguments: XB ≥ 50 or Y ≥ 150
drops the command, and NB is clamped to 50 − XB (a resulting NB = 0 drops it).
All three argument bytes are consumed before any drop, so the stream stays aligned.
That bounds the highest byte it can touch at $BA65 — safely inside the image and
nowhere near the $BFE0 registers. Writes are a plain store, not read-modify-write,
so the odd sub-columns of each touched byte are cleared.
The CPU builder gpu_hdotline takes ordinary half-res pixel coordinates and
auto-aligns them to this encoding.
DOT_POLYGON ($4C), POLYGON ($4D), POLYGON16 ($4E)#A closed figure around a defined centre, rotated and scaled by the GPU, and always clipped. One command draws a whole outline:
$4C | $4D | $4E CX.16, CY.16, ANGLE, SCALE, N, dx0,dy0, … , dxN-1,dyN-1
| Field | Meaning |
|---|---|
CX, CY |
The centre, signed 16-bit, and it may lie off screen. Half-res units (the 200 × 150 grid extended to signed 16 — the same encoding gpu_dotline_clip already takes) for $4C and $4D; full-res (400 × 300) for $4E. |
ANGLE |
Rotation about that centre, one byte of brad: 256 brad = a full turn, 128 = half. Same convention and the same direction as API_SIN/API_COS ($FF7B/$FF7E). |
SCALE |
Unsigned Q0.7: 128 = 1:1, 64 = half. Shrinking only — the GPU clamps anything above 128 down to it, and treats 0 as "draw nothing". |
N |
Bit 7 is the CLOSED/OPEN flag; bits 0–6 are the vertex count K (0–127). Bit 7 = 0 (default): CLOSED — the last vertex joins the first, so K vertices are K segments (today's behaviour, unchanged for every K an existing caller sends). Bit 7 = 1: OPEN — K vertices are K-1 segments, a true polyline through the same rotate+scale matrix. |
dx, dy |
Signed byte offsets from the centre, unrotated and unscaled — the raw shape. Half-res for $4C/$4D, full-res for $4E. |
CLOSED is a figure around a point; OPEN is a rotated, scaled polyline — N's top bit
picks which. No _CLIP suffix, because the anchor is signed-16 and may sit off screen, so
the family is always clipped either way. For a path with no rotate/scale matrix,
or more than 127 vertices, use the chains ($45/$4A/$4B) instead.
It generalises DOT_CIRCLE ($48) (already anchored at its centre, but with a byte
anchor that must be on screen) — $48 is untouched.
$4E unlocks#A CPU-side quarter-square multiply keeps a transformed vertex at half-res precision at
best (doubling only ever yields even coordinates). The GPU instead interprets the
same ±127 offset as full-res directly: a figure up to 254 px across (the limit —
±128 no longer fits a signed byte), each rotated offset rounding to a whole full-res
pixel — full-res figures cannot be produced any other way. $4E costs roughly twice
$4C's pixels (full resolution vs. every-other-row) plus LINE16's +5–8% for 16-bit
endpoints.
Measured in py65 on a 12-vertex, radius-48-half-res figure. "Transform + clip" is what moved onto the GPU versus the same shape pre-transformed and sent as a chain.
| wholly on screen | crossing an edge | rejected | transform + clip | |
|---|---|---|---|---|
$4C DOT_POLYGON |
26 923 (11.3% of a frame) | 21 271 | 481 | 8 000 (666/vertex) |
$4D POLYGON |
39 513 (16.6%) | 26 040 | 483 | 8 310 (692/vertex) |
$4E POLYGON16 |
41 387 (17.4%) | 28 934 | 483 | 7 481 (623/vertex) |
A frame is 237 404 cycles; cost scales linearly with vertex count, so level-of-detail
(fewer vertices for a small figure) is entirely CPU1's to pull — there is no shape table
or SHAPE_ID on the GPU. An off-screen figure costs ~480 cycles flat, about a sixtieth
of drawing it. This is a transfer, not a speed-up — the same algorithm costs the same
on either 65C02; what moves is which processor is idle when it runs.
The GPU folds rotation and scale into one matrix built once per figure (so the scale rides free inside the rotation's products) and clips with a Cohen–Sutherland outcode test per segment, sharing the chain's joint saving between two ends that are both inside. It clips, never clamps — a clamped corner would bend every segment touching it, so true signed-16 geometry is kept until the cut, a division-free bisection (sixteenths of a unit, ≤4 passes) snapped onto the crossed boundary.
$80 (−128) is clamped to 127.SCALE shrinks only. Above 128 clamped to 128; 0 draws nothing.K = 0 draws nothing and consumes no vertex bytes, CLOSED or OPEN. CLOSED
K = 1/K = 2 are degenerate but harmless (a dot; one edge drawn twice). OPEN
K = 1 draws nothing — one vertex has no segment.2K argument bytes, so a rejected
figure never leaves the dispatch loop reading vertex data as opcodes.API_SIN#The GPU carries its own sine/quarter-square tables (1152 B at $F300/$F600–$F9FF —
the two processors can't see each other's ROM). The quarter-square multiplier
(a·b = f(a+b) − f(|a−b|), f(x) = ⌊x²/4⌋) costs ~60 cycles/product against ~130 for
bit-serial. The sine table is scaled 0–128, not the CPU OS's 0–127 Q0.7: that makes
ANGLE=0, SCALE=128 an exact identity and a quarter turn exact, at the cost of a
1-unit-in-128 disagreement with API_SIN/API_COS on the same angle — never in
direction (both turn +X toward +Y, clockwise on screen).
roms/test_polygon.py: exact pixel-set equality against the
identity case ($4C==$45, $4D==$4A, $4E==$4B), an independent Liang–Barsky clip for
edge-crossing figures, an N=1 probe of the rotate/scale matrix directly, and rotation
direction against API_SIN/API_COS.
CIRCLE16 ($4F) — the circle with an anchor that may be off screen#$4F CX_LSB, CX_MSB, CY_LSB, CY_MSB, R (6 bytes)
The circle counterpart of the polygon family: a solid outline around a signed-16
full-res centre that is allowed to sit off screen, and therefore always clipped.
R is a radius in full-res pixels, 0–255; R = 0 draws nothing. Builder:
gpu_circle16 ($FFDB).
$48 DOT_CIRCLE is untouched and is not superseded — it remains the dotted, half-res,
on-screen-only call, at about half the cycles ($4F runs the same midpoint walk at
full resolution, which is what makes its outline solid).
No ANGLE, no SCALE, no half-res twin. ANGLE does nothing to a circle (the
polygon family's rotate/scale matrix has nothing to compose here); SCALE would ride in
no existing product (one extra multiply + stream byte to save the caller one it already
does per object — a composite object still works from one struct, only the anchor
shared); and a half-res form would cost the same 6 stream bytes and cycles as this one,
buying only a radius range past 255 px in exchange for an even-only centre. The gap that
would earn an opcode is the dotted twin, DOT_CIRCLE16 — see
MAD65_future_features.md.
Solid and clipped exactly. 8-connected, 1 px thick, same octant walk as $48 but at
full resolution ($48 doubles the half-res walk instead, hence dotted). Clipping is
decided once from the bounding box:
| Bounding box | What happens | Cost |
|---|---|---|
| misses the screen | nothing drawn | 200 cycles |
| wholly inside | fast path, no per-point range tests | see below |
| straddles an edge | per-point tests every column/row | +8–15% |
Clipped per point, not geometrically — no intersection to compute — so $4F is the
one clipped primitive whose output is exactly predictable:
roms/test_circle16.py holds it to
midpoint_walk(cx,cy,r) ∩ screen with no tolerance.
| Radius (full-res px) | wholly on screen | half off an edge |
|---|---|---|
| 10 | 6 426 (2.7% of a frame) | 7 336 |
| 50 | 27 571 (11.6%) | 31 626 |
| 150 | 105 039 (44.2%) | 93 074 |
Linear in R (every point is a read-modify-write). A 100 px-radius circle is a quarter
of a frame and should be a deliberate choice; $48 at the same visible size costs about
half.
The GPU drawing routines do not validate coordinates. Out-of-range values are not caught, clipped or rejected — the GPU computes whatever address the formula gives and writes there. CPU1 must ensure every coordinate is in range before it reaches the instruction list.
| Instruction family | Parameter | Legal range | Notes |
|---|---|---|---|
PIXEL |
X, Y | 0–399 / 0–299 (16-bit) | Full-resolution |
LINE, DOT_LINE, DOT_LINES, DOT_CIRCLE |
X, Y | 0–199 / 0–149 | Half-res, doubled internally |
DOT_PIXEL, DOT_PIXELS |
X, Y | 0–199 / 0–149 | Same space as lines |
TEXT, TEXT_BG, TILE, TILE_BG |
column, row | 0–49 / 0–36 | Tile grid units (50 × 37 cells) |
SPRITE, COMB_SPRITE |
X, Y | signed 16-bit | Top-left pixel, (0,0) = screen origin; GPU clips all four edges |
COMB_SPRITE |
GAP | 0–15 | Masked with $0F; the high nibble is reserved and ignored |
DOT_CIRCLE |
CX, CY | 0–199 / 0–149 | Centre must be on-screen (early-out if violated) |
HDOT_LINE |
XB, Y, NB | 0–49 / 0–149 / 1–50 | Self-clipped — see above |
DOT_POLYGON, POLYGON |
CX, CY | signed 16-bit, half-res units | May be off screen — the GPU clips every edge |
POLYGON16 |
CX, CY | signed 16-bit, full-res units | May be off screen — the GPU clips every edge |
$4C–$4E |
dx, dy | −127…+127 (signed bytes) | −128 is clamped to 127; the figure is at most 254 units across |
$4C–$4E |
SCALE | 0–128 (Q0.7, 128 = 1:1) | Shrinking only; above 128 clamped to 128, 0 draws nothing |
CIRCLE16 |
CX, CY | signed 16-bit, full-res units | May be off screen — the GPU clips every point |
CIRCLE16 |
R | 0–255, full-res pixels | R = 0 draws nothing |
The address formula for pixel and line instructions is:
VRAM address = $8000 + Y_full × 50 + X_full / 8
VRAM-image ends at $BFDF. Immediately above it sit write-only hardware registers:
LED_GPU_REG at $BFC0 and VIDEO_REG at $BFE0 (BG_REG bit 0, BLINDER bit 1,
COPY_DIS bit 2, SHADOW_MODE bit 3). If Y_full ≥ 326 the computed address reaches
$BFC0 and a pixel write latches random bits into one of them — BLINDER blanks the
screen instantly, BG_REG switches the displayed layer, COPY_DIS kills the background
copy, and clearing SHADOW_MODE un-maps the shadow RAM the GPU is executing from.
All persist until the next valid VIDEO_REG write, and the last one does not survive to
make one. For half-res instructions the threshold is Y ≥ 163. Values above the legal
maximum but below the threshold write into the unused VRAM tail — invisible, but still
wrong.
The two register ranges are decoded separately (VIDEO_REG needs A7 & A6 & A5), so
GPU_LED is safe to issue at any time, including during rendering. See
MAD65_address_decoder.md §3/§4.
Bounds checking costs cycles on every draw call for coordinates that are almost always
valid; validation belongs in CPU1's scene-composition logic, where it runs once a frame
instead of once a pixel. DOT_CIRCLE's centre check, SPRITE's clipping, HDOT_LINE's
row-end clamp and the polygon family's clip are exceptions only because those routines
need the range information to drive their own inner loops (bounding box, clip region,
byte count) — a per-segment clip in particular runs once per SEGMENT, not once per pixel,
so it costs the same either processor and only changes which one is idle. Pure drawing
instructions like PIXEL/LINE have no such internal need and stay unchecked.
RECT_BG ($31), RECT_BG_RLE ($32), LOAD_RLE ($33)#Three opcodes added together because they are one job. None of them is drawing: they
move bytes from the PPRAM command stream into GPU memory, which is why they live beside
LOAD in the $30 block and why they are dispatched from dispatch_rare rather than
the fast chain — a background load is a handful of commands a frame across ~17 frames,
not a per-frame draw. The fast chain keeps its last free slot for a future drawing
opcode, where 3 cycles a command would actually matter.
LOAD could not do#$30 is a linear 256-byte copy to a page-aligned destination, and a page straddles 5.12
screen rows — structurally row-agnostic, unable to express an arbitrary rectangle or a
comb. $31 takes a byte column and a row and walks the destination by 50·(GAP+1).
ctrl $01–$7F the next ctrl bytes are literals
ctrl $80–$FF repeat the ONE byte that follows, (ctrl − $80) + 2 times ($80 = 2 … $FF = 129)
ctrl $00 never emitted
A repeat run is 2 stream bytes for up to 129 output bytes; a literal run costs 1 byte of overhead per 127, so the worst case on incompressible data is +1/127 ≈ +0.8 %. Measured on this repository's own 1-bit art it returns 1.4×–2.3×, which is roughly three quarters of what a general-purpose compressor finds, for about a hundred bytes of decoder.
There is no end marker and no length header. The decoder is driven by an output
count the consumer already knows — $32 consumes WB·H, $33 takes its count as an
argument. That keeps the decoder state to three bytes, and an over-long final run is
harmless.
The producer's obligation, and it is NOT checked at runtime. A stream too short for its output count is read straight past its end and into the next command, and the dispatch loop desynchronises — arbitrary bytes then execute as opcodes. For
$32that means every band must encode exactlyWB·Hbytes, so a picture whose height is not a multiple ofHneeds its last band padded.roms/rle.pydoes this. Checking it here would cost a byte budget in every command to catch a build-time mistake.
The background layer is transport-bound, not blit-bound (~3.5 pages/frame through the
ping-pong window), so compressed bytes must cross PPRAM and expand here — decompressing
on CPU1 would send raw bytes through the one bottleneck that matters and win nothing. The
mirror case — a payload sinking in CPU1's own RAM — never touches PPRAM, so CPU1 carries
its own decoder (api_decrunch). Compress at the source, decompress at the sink.
The two decoders are deliberately different codecs. This plane is write-only, so a
decoder here can never look back at what it has written and has to be stateless — RLE.
CPU1's sink is ordinary RAM, so api_decrunch runs LZ (roms/lz.py),
which copies from its own output and beats RLE on everything CPU1 loads (the demo's VGM
song 6.3×, where RLE gets 1.00×). Until 2026-09-11 both sides ran this RLE.
Background writes go through DRAW_BASE_HI = $C0, so the address is $C000 + row·50 + XB
in a 16-bit pointer. The framebuffer's last byte is $FA97. Past $FFFF the pointer
wraps to $0000 — zero page and the stack, not a harmless corner of VRAM. The unused tail
$FA98–$FFFF is 1384 bytes, which is 27 rows at GAP = 0 and exactly one row at
GAP = 15: the comb multiplies the row step by up to 16 and eats the margin.
Every row is therefore bounds-tested and an out-of-range row is sunk — its payload is read and discarded. Sinking rather than stopping is the point: with an RLE payload you cannot skip forward without decoding, and stopping early would leave the rest of the stream in PPRAM for the dispatch loop to execute.
The horizontal edge is deliberately not checked. XB + WB > 50 smears one row into
the start of the next — cosmetic, confined to the background plane, and a loader contract
rle.py checks offline.
About 45 cycles a payload byte (the source vector, the fetch and the store). During a
background load CPU1 can push ~900 payload bytes a frame, so roughly 40 k cycles — about
17 % of the frame, on frames where the game is loading rather than playing. A worked
example with real numbers is in carts/rect_bg_test/.
Any instruction that writes VRAM-background must be sent twice, on two
consecutive frames: CLEAR_BG ($20), TEXT_BG ($61), TILE_BG ($71), and
LOAD ($30) to a $C0–$FF page.
Why: VRAM is double-buffered in hardware, and a flip-flop swaps which physical chip is "the background" on every VSYNC. The GPU never knows which one it holds, so a single write lands in only one buffer; the next frame the hardware swaps to the other, which still holds the old content — the change flickers at 30 Hz or vanishes. Writing the same data on two consecutive frames reaches both buffers, after which it is stable.
Image-layer instructions do not need this. The image buffer is rebuilt every frame, so a one-frame write is correct there. Hardware copies the background onto the image each frame automatically, so the background only needs redrawing when its content actually changes — but when it does, across two frames.
The CPU OS automates this with an argument-capture replay engine, so a game issues a
background write once — see MAD65_CPU_OS.md.
$0700–$0EFF ($0600–$06FF is
SPR_HEIGHT, the 4th sprite-table page).$00 is the end-of-string sentinel — never define or use it as a drawable tile.TILE_BANK + tile_index × 8 = data address.Example (letter 'A'):
00011100 = $1C 01111111 = $7F
00110110 = $36 01100011 = $63
01100011 = $63 01100011 = $63
01100011 = $63 00000000 = $00
TEXT)#$20–$7F, 8 bytes each = 768 B, in ROM at $FC00–$FEFF.font_data + (char − $20) × 8.$60, $61, $64, $73, $77 are repurposed as symbols (copyright
sign, arrow glyphs), not lowercase letters.roms/fonts.s.TEXT/TEXT_BG take a scroll argument (0–7). The whole rendered row buffer is
treated as one wide binary number and shifted left by scroll pixels, bits
carrying right-to-left between adjacent cells, so the shift is pixel-continuous across
the string.
The last character of the string is never written to VRAM — it is the scroll
buffer. Its high bits do carry into the rightmost visible cell during the shift, so
the next character scrolls in smoothly from the right. A left-scrolling marquee is
therefore trivial: hold a window of N+1 characters (N visible + 1 guard), increment
scroll 0→7 one pixel per frame, and when it wraps advance the string pointer by one
character. Static text must end with a space.
The visible draw count TEXT_LEN is total − 1, capped at the 50-cell row buffer and
clipped at the right screen edge (50 − column). scroll = 0 skips the shift entirely
(static-text fast path). A 51-byte string fills all 50 columns; longer strings clip,
never overflowing onto the next row.
⚠ A scrolling string must stop one cell short of the right edge. Phase 0 stores at most 50 cells, so a 51-character string at column 0 gives 50 visible cells with the guard pushed out of the buffer —
TEXT_LEN == stored, no+1, and zeros shift into column 49. The text still fills the width, but each new character pops into existence instead of sliding in, once every 8 frames. A marquee therefore wants 50 characters (49 visible + guard), or more generallyvisible ≤ 49 − column. Filling all 50 columns and smooth-scrolling are mutually exclusive on this path.This is a
TEXT/TEXT_BGrule only.VTEXT/VTEXT_BGhave no scroll guard — their scroll is an address offset rather than a bit-shift chain and every character is drawn — so a vertical marquee may use the full 37-cell width (seecarts/tate_test).
VTEXT)#VTEXT ($62) / VTEXT_BG ($63) draw text for a monitor turned 90° clockwise
(madsim's F12 mirrors the turn on screen). They are the only orientation-aware
routines: a quarter turn maps the 400×300 framebuffer onto itself, and everything else
already works in framebuffer coordinates, so sprites, lines, tiles and LOADs just need
pre-rotated assets, no code change. See MAD65_future_features.md.
The grid is 37 × 50 — the exact transpose of TEXT's 50 × 37, given in the same
(X, Y) order and directions as every other cell-grid call:
| argument | range | on screen | in the framebuffer |
|---|---|---|---|
X — char cell |
0–36 | left → right, 37 cells (296 of 300 px, 2 px margin) | fb-y = 290 − 8×X (decreasing) |
Y — line |
0–49 | top → bottom, 50 lines (400/8 exact) | byte column Y (fb-x) |
scroll |
0–7 | shifts the string LEFT, 1 px at a time — same direction as TEXT |
adds 50 × scroll to the address |
Screen position is 2 + 8×X − scroll px from the left edge — the identical marquee
recipe TEXT uses, but here it's free: along fb-y a pixel is a whole VRAM row, so
scroll is an address offset, not a bit-shift chain, and VTEXT draws every
character (no hidden last cell like TEXT). roms/test_vtext.py pins the shared
direction. To scroll the other way, count scroll down 7→0 instead — it's an offset,
not a direction (see carts/tate_test/).
Cells 37/38 exist but can't be named (X > 36 draws nothing) — reached only by a
string longer than its line, they're what lets a character enter from the right edge
one pixel column at a time. Two guard cells, not one, because a cell is 8 px wide and
scroll spans 0–7. Clipping is per glyph row, right edge only:
$7800 PPRAM window); each row is tested with one BMI.X ≤ 36 the largest reachable offset is 15249,
inside the 16352-byte VRAM window's invisible tail (video circuit stops reading at
15000, register overlay starts at 16320) — a glyph scrolled off the left simply
disappears into scratch memory for free. The X range check is what makes this
safe, and gpu_os.s asserts that bound. X > 36 or Y > 49 draws nothing (string
still consumed, so the dispatcher stays in sync).Glyphs come from font_data_v at $F000, generated at build time from
roms/fonts.s by roms/gen_font_v.py —
never hand-edited. It sits exactly $0C00 below font_data (linker-asserted), letting
draw_vtext_core reuse font_glyph_lo[] with only the high byte changed.
Builders gpu_vtext/gpu_vtext_bg ($FFB4/$FFB7) pass the cell through
unclamped (clamping a signed coordinate would destroy $FE/$FF); the GPU clips it.
$0300, 4 pages, one parameter per page for fast indexed
access. Up to 256 sprites, indexed by SPR_ID.roms/sprite_zero.s, $FF4A). boot_main installs it so
id 0 is usable immediately; a developer may overwrite the slot at any time.(0,0) = screen origin — the same zero as PIXEL and the clip builders.spr_setup does a coarse 16-bit reject per axis first, which also guards the
signed-byte internal column count against aliasing for far-off coordinates. Visible
range: X in −(8·W)+1 … 399, Y in −H+1 … 299.$1000 onwards, uploaded via LOAD.| Page | Address range | Content |
|---|---|---|
| Page 0 | $0300–$03FF |
SPR_TYPE |
| Page 1 | $0400–$04FF |
SPR_PTR_LSB |
| Page 2 | $0500–$05FF |
SPR_PTR_MSB |
| Page 3 | $0600–$06FF |
SPR_HEIGHT |
Width and overlay come from the type byte: low nibble = width in bytes W
(1 = 8 px, 2 = 16 px, 4 = 32 px, 8 = 64 px); bit 4 ($10) set → no
overlay plane, clear → overlay present.
Height is independent, in SPR_HEIGHT ($0600 + SPR_ID): any value 1–255 rows.
It is not derived from width, so non-square sprites are allowed (an 8×64 column, a
32×8 bar). Height 0 marks an undefined slot and renders nothing — boot zeroes the
table, so unused ids are inert. The blitter costs the same per row regardless of
height, so tall sprites cost proportionally more only because they have more rows.
| Type | $01 |
$11 |
$02 |
$12 |
$04 |
$14 |
$08 |
$18 |
|---|---|---|---|---|---|---|---|---|
| Width | 8 px | 8 px | 16 px | 16 px | 32 px | 32 px | 64 px | 64 px |
| Overlay | yes | no | yes | no | yes | no | yes | no |
64 px (
W=8) is the widest the type byte can encode:W=16(128 px) would be$10, which collides with the overlay-absent bit. 128 px would need a re-encoded type byte.
1 bit per pixel, packed 8 px/byte, MSB = leftmost — same bit order as tiles and glyphs. Stored interleaved by row:
W bitmap bytes then W overlay bytes (stride 2W).W bitmap bytes (stride W).| Plane | Bit = 1 | Bit = 0 | Operation |
|---|---|---|---|
| Bitmap | white | transparent | OR (screen \| bitmap) |
| Overlay | black | no change | AND NOT (& ~overlay) |
The composite is a single read-modify-write per byte:
screen = (screen OR bitmap) AND NOT(overlay)
The bitmap draws the white silhouette; the overlay punches black detail on top. Because both planes treat a 0 bit as "leave alone", a sprite shifted to a non-byte-aligned column needs no boundary masking — the zero bits shifted in on the spill side are inert, and partially-visible edge bytes come out correct without per-bit clipping.
COMB_SPRITE ($51)#$51 takes SPRITE's five arguments unchanged in meaning and order, plus a sixth
byte GAP (0–15, masked with $0F). With D = GAP+1, source row r is blitted to
screen row Y + r·D; the GAP rows in between are never written, so whatever the
hardware copied under the scene stays visible. On-screen height is
(ROWS−1)·D + 1, i.e. the same sprite data covers up to 16× its own height.
SPR_TYPE, same SPR_HEIGHT, same pixel
data. One sprite can be drawn solid with $50 and as a ghost with $51 in the same
frame — ghosting is a transient state (hit flash, phasing, spawn), not a property
of the artwork, which is exactly why it is an opcode and not a type-byte bit.SPR_STRIDE still steps a whole interleaved row, so row r reads from the right
place). An overlay would punch the background to black on drawn rows while the gap
rows keep it, which reads as shredding the background rather than as transparency.
The cost of this is that a comb sprite can never occlude what is behind it.GAP stretches and thins: at GAP=3 the object is 4× taller with
only 25 % of its rows drawn, which reads as coming apart — good for an explosion or a
dissolve, not a substitute for a real zoom.Y a
multiple of D and it moves in D-pixel steps instead of strobing. Two comb sprites
at Y and Y+1 (same GAP=1) interleave into one solid image — the phase is the
Y parity, so there is no separate phase argument.GAP — and falls as the object
grows past the screen edges and rows clip away. On-screen height is free; source
rows and width are what you pay for.Budgeting it (measured on carts/comb_test, unit = one 64 px-wide slot, one source
row; a W-px-wide, H-row sprite is ceil(W/64) × H of them): ~330 cycles/slot-row
byte-aligned, ~570 at the worst bit offset (spr_setup picks the cheaper shift
direction, so it's a tent peaking at offset 4, not a ramp). A 256×83 source filling the
screen at GAP=1 runs 46–79% of a frame — a comb sprite that wide is a scene, better
placed in the background layer unless it's the only thing moving; a 128×41 one at
12–20% is an ordinary moving object, and GAP then adds height for free (same 12% at
161 px tall, GAP=3).
Clipping is spr_comb_vclip, which differs from $50's vertical clip in exactly two
places, both of which are where the bugs would be:
top_skip = −Y = D·q + r, the first
source row drawn is ceil(top_skip/D) and the first visible screen row is
(r == 0) ? 0 : D − r — anywhere in 0…D−1. $50's clip only ever produces 0.ceil((300 − fr)/D), not 300 − fr. D ≥ 2 bounds it at
150, which keeps the row count 8-bit and proves the last written row is ≤ 299.
Dropping the /D would run the row loop past the image end at $BA98 into the video
registers at $BFE0 — the same hazard class as the old lc_vline bottom-clip bug.
GAP=0 (D=1) deliberately keeps $50's clip, because there alone the count
reaches 300.The division that both of these need (div_small, 16 ÷ 1–16) runs at most twice per
sprite and never per row; the top-clip call only happens when the sprite actually hangs
off the top edge. Tests: roms/test_comb_sprite.py,
which checks every residue class of top_skip mod D and asserts on every case that
nothing was written at or past $BA98. Working example:
carts/comb_test/ — a ghost planet drifting over background
captions that stay readable straight through it, with GAP on the joystick and an
end-to-end harness that measures the see-through property rather than eyeballing it.
The GPU never returns from its interrupt — VSYNC is the frame trigger, so frame_isr
resets SP on entry and ends in WAI.
VSYNC interrupt fires
Reset SP ← $FF (discard the hardware-pushed frame; there is no caller)
If FRAME_BUSY: ← the previous frame overran; see the policy below
OVERRUN_CNT += 1 (saturating)
FRAME_BUSY ← 0
CLI (a late VSYNC must not be blocked)
If NO_CPU_COUNT >= $40: ← diagnostic mode check, before anything else
PPWP ← DIAG_LIST ($FA00)
goto dispatch_loop (PPRAM status check skipped entirely)
PPWP ← PPRAM ($7800) ← normal mode
Read status byte via PPWP (PPWP advances to $7801)
If status ≠ CPU_READY:
NO_CPU_COUNT += 1
status ← GPU_READY
WAI (nothing rendered this frame — a visible blink)
NO_CPU_COUNT ← 0
If BOOT_BLINDED: ← one-shot: first real frame drops boot's BLINDER
clear BLINDER, clear BOOT_BLINDED
status ← GPU_WORKING
FRAME_BUSY ← $80 (armed: a VSYNC before the end of the list is an overrun)
(PPWP already at $7801 = first opcode)
dispatch_loop:
Read opcode via PPWP, increment PPWP
If opcode = $00 (WAI):
FRAME_BUSY ← 0 (finished inside its VSYNC period)
status ← GPU_READY
WAI
Dispatch to the opcode's routine
Routine reads its own arguments via PPWP and advances it
goto dispatch_loop
Convention: whenever a byte is read via
PPWP, incrementPPWPimmediately, so it always points at the next unread byte. Opcode routines end withJMP dispatch_loop, notRTS— this keeps the call depth flat.Dispatch is a linear
CMPchain ordered by call frequency, hottest first (dotted geometry → pixel/sprite → text/tile → solid line → polygon → circle →LOAD→WAI→ rare control ops). Cost is30 + N×4cycles for the Nth entry. AllBEQtargets are 3-byteJMPtrampolines listed in the same order as the comparisons, becauseBEQonly reaches ±128 bytes; the assembler errors if one falls out of range.The chain is capped at 29 entries. The first
BEQhas to reach the first trampoline, which is 117 bytes ahead at 29 entries and out of range at 30. The polygon family took the last three slots andCIRCLE16a fourth, soCOPY_DIS_ON/_OFF,BG_REG_ON/_OFFandGPU_LEDmoved intodispatch_rare— aBNE+JMPtail with no reach constraint, reached by falling off the end of the chain. All five are written at most once per state change, never inside a per-frame loop, so the extraJMPcosts nothing that matters. The chain currently holds 28. A new opcode goes intodispatch_rare, or earns its way up by displacing something.
FRAME_BUSY is set just before the dispatch loop is entered and cleared when the list's
WAI is reached. Finding it set on entry to frame_isr means a VSYNC arrived while
the previous frame was still drawing: that frame overran its budget.
The policy is "the new frame wins" — there is no choice about it. The ping-pong
hardware has just swapped the GPU's 2 kB window, so the abandoned dispatch's command
list is gone; resuming would splice one list into the middle of another (HUD text
executed as opcodes: ' ' is CLEAR_BG, '0' is LOAD, and further on lie VIDEO_REG
and BLINDER). The SP reset at the top of frame_isr has already discarded the old
frame, and it never returns.
A viewer sees one partially drawn image, for exactly one frame — the hardware
background copy re-lays every frame, so a truncated image never persists: consistent
overrun produces torn frames, not accumulating wreckage, and recovers the moment load
drops. CPU1 is not told (no handshake on GPU_READY); the GPU only counts:
OVERRUN_CNT saturates at 255, read by nobody but a bench or debugger. "Finish the
frame you started, drop the one you missed" is not available — it would need the
in-flight list to survive the VSYNC, and the swap is unconditional hardware.
$C000–$FFFF, 16 kB)#$C000–$C002 JMP stub — RESET → boot_main
$C003–$C005 JMP stub — IRQ → frame_isr
$C006–$???? Boot procedure (boot_main)
$????–$EFFF Frame handler, dispatch loop, op_* routines, helpers
$F000–$F2FF Rotated font (TATE) 768 B 96 glyphs × 8 B (FONT_V segment)
$F300–$F37F Polygon |sin| table 128 B pg_sinmag, 0-128 scale, brad.
128 entries, so 128-alignment is
all an indexed load needs
$F380–$F3DD Vertical text tables 94 B vtext_col_lo/hi + vtext_scr_lo/hi
$F3DE–$F3FF Reserve 34 B filled $EA
$F400–$F5FF VRAM row-start tables 2×256 B row_lo ($F400) / row_hi ($F500)
$F600–$F9FF Polygon quarter-square mult 4×256 B pg_qs_lo/hi ($F600/$F700),
pg_qr_lo/hi ($F800/$F900) = the
same table + 64 (rounding term)
$FA00–$FAFF Diagnostic instruction list 256 B 82 B used (gpu_diag_list.s)
$FB00–$FBBF Font glyph pointer tables 2×96 B FONT_GLYPH_LO ($FB00) / _HI ($FB60)
$FBC0–$FBFF Reserve 64 B filled $EA
$FC00–$FEFF Font definitions 768 B 96 glyphs ($20–$7F) × 8 B
$FF00–$FF49 Text row-start tables 2×37 B TEXT_ROW_LO ($FF00) / _HI ($FF25)
$FF4A–$FFB1 Default sprite 0 104 B 32×26, no overlay (SPRITE0 segment)
$FFB2–$FFF9 Reserve 72 B filled $EA
$FFFA–$FFFB NMI vector (unused — points to irq_stub)
$FFFC–$FFFD RESET vector → $C000
$FFFE–$FFFF IRQ/BRK vector → $C003
Segment starts are pinned in roms/gpu.cfg; the CODE end moves
with every build (see gpu_os.map). CODE owns one contiguous run, $C000–$EFFF,
and currently ends at $E9AA — 1621 bytes free. Data segments are page-aligned only
where an indexed load can actually cross a page (row_lo/row_hi, the four
quarter-square pages); anything reached through a pointer or bounded inside a page is
packed tight instead. Four small islands remain that CODE cannot reach — 344 B total,
named as reserve in gpu.cfg rather than left looking like an oversight.
Two placements are load-bearing, and a linker assert fails the build if either is disturbed:
FONT sits exactly $0C00 above FONT_V. draw_vtext_core reuses the horizontal
font's glyph-pointer table and only shifts the high byte. That pair is also what fixes
the $F000 ceiling on CODE, so the assert at the end of the CODE segment says so in
plain language — ld65 does fail on an overrun, but it fails with "Segment 'FONT_V' start
address is too low", which reads as a font problem rather than "the code got too big".DIAG_LIST is at $FA00. gpu_os.s carries the address as a constant, because
gpu_diag_list.s is its own translation unit with no label to import. A second assert
keeps the list inside its page: growing past $FB00 would silently overwrite the glyph
base table and corrupt every TEXT the GPU draws. The list has 256 bytes and uses 82.The two JMP stubs give the hardware vectors fixed targets while the real handlers live
anywhere in the image. row_lo/row_hi, the quarter-square pages and the font are
page-aligned deliberately — see Implementation invariants.
If frame_isr counts 64 consecutive frames without CPU_READY (~1 s),
NO_CPU_COUNT reaches $40 and the GPU enters diagnostic mode. From then on
frame_isr points PPWP at DIAG_LIST ($FA00) instead of PPRAM and jumps straight
into the dispatch loop, which is oblivious to where its opcodes come from.
NO_CPU_COUNT is never reset in diag mode — the CPU is not expected to recover; a
missing CPU signal after boot is fatal.
The list (roms/gpu_diag_list.s) uses the same opcode format
as PPRAM but has no status byte — opcodes begin at byte 0 — and must fit 512 B and
end with WAI. Current content: BLINDER_OFF (the diagnostic screen reveals itself,
since boot leaves the screen blinded), sprite 0, and three TEXT lines —
MAD-65 PROTOTYPE V0.9 / CPU NOT READY / GPU IN DIAGNOSTIC MODE. It doubles as a
graphics self-test with no CPU involvement; widening it (border, full ASCII sweep,
VIDEO_REG flag cycling) is a matter of adding opcodes to that file.
SEI + CLD disable interrupts, force binary mode
VIDEO_REG ← BLINDER set, COPY_DIS + BG_REG clear (still SHADOW_MODE=0 boot mode)
LED_GPU_REG ← $01 stage 0: boot started
PPRAM[$7800] ← GPU_BOOTING ($B0) + ZP mirror — the FIRST PPRAM write of boot,
deliberately BEFORE the ~1-frame shadow copy
NO_CPU_COUNT ← 0 (the ZP clear below re-zeroes it anyway)
init stack (S ← $FF)
copy ROM $C000–$FFFF → shadow code-RAM boot mode: read hits ROM, write hits shadow.
SHADOW_MODE stays 0 until just before the first
$C000–$FFFF write (the VRAM-background clear)
init zero page wipes GPU_STATUS / VIDEO_REG_SHADOW — intentional
PPRAM[$7800] ← GPU_BOOTING ($B0) + ZP mirror, AGAIN: re-establishes the wiped mirror
and re-asserts the byte on the chip owned by now
init RAM ($0100–$77FF)
LED_GPU_REG ← $03 stage 1: RAM cleared
install default sprite 0 into the definition table
VIDEO_REG ← BLINDER | SHADOW_MODE enter run mode: code now runs from shadow RAM and
$C000–$FFFF writes reach VRAM-background. Must land
AFTER the ZP clear (which wiped VIDEO_REG_SHADOW)
and BEFORE the first $C000+ write below
clear VRAM-background × BG_CLEAR_PASSES back-to-back loop, NO waits between
re-assert BLINDER (local invariant: blind while stamping the bg)
draw build-date stamp → the background bank owned right now
wait one VSYNC bg bank ownership swaps
draw build-date stamp → the other bank
wait one VSYNC let the bg→image copy populate both image banks
LED_GPU_REG ← $07 stage 2: VRAM cleared and stamped
BOOT_BLINDED ← 1 arm the reveal — the screen stays DARK for now
PPRAM[$7800] ← GPU_READY ($B2) + ZP mirror
LED_GPU_REG ← $0F stage 3: GPU ready
CLI enable interrupts
WAI
The screen is not unblanked at the end of boot. Revealing here would flash the bare
datestamp on an empty background for every frame the CPU takes to come up. Instead boot
arms BOOT_BLINDED, and the reveal happens on the first CPU_READY frame, inside
frame_isr, just before the first PPRAM list is drawn. It is a one-shot, so a game that
later drives BLINDER itself is never overridden. The diagnostic screen reveals itself
via DIAG_LIST's own BLINDER_OFF.
Why the background clear is a tight loop. Only the background is cleared —
the video circuit copies background → image every frame even while BLINDER is set, so
a clean background blanks the display on its own. The clear must fully zero both
double-buffered halves; one pass takes ~0.6 of a frame, so must not be phase-locked
to VSYNC ("clear; wait exactly one frame; repeat" made every pass start at the same
phase, leaving one buffer with a permanently orphaned region — a 30 Hz flicker after
reset, since SRAM retains pre-reset content). The fix is a back-to-back loop with no
waits: ownership still flips once per VSYNC, but the clears land at every sub-frame
phase, so each buffer gets at least one complete pass. BG_CLEAR_PASSES is 12
(~0.12 s); measured, both buffers are zero by frame ~2–3 — wide margin.
The build-date stamp is drawn after the clear loop (so no pass can erase it) and
before the reveal. It lives in the background, so the free hardware copy re-paints it
every frame and it survives into whatever the CPU draws, as long as the program leaves
the top-right corner alone. It is tiny (14 cells × 8 rows) and finishes well inside one
bank's ownership window, so unlike the full clear it never straddles a swap — hence the
simple draw / wait one VSYNC / draw again. BUILD_DATE is a "YYMMDD" literal the
Makefile regenerates into build_date.inc before every assembly.
GPU_BOOTING is written twice)#PPRAM survives a warm reset, so after a reset both ping-pong chips still hold the
previous run's status bytes — one ends with CPU_READY, the other with GPU_READY.
Each core can only write the chip it currently owns (they swap every VSYNC), so a single
BOOTING write sanitises only one of the two. A surviving stale byte false-triggers
either direction: stale GPU_READY lets the CPU's boot-tail wait pass while the GPU is
still clearing RAM (losing the CPU's first command lists, including one-shot LOADs);
stale CPU_READY makes the GPU's first frame_isr execute a leftover pre-reset list.
The defence: both cores announce BOOTING immediately at boot entry, before their slow ~1-frame shadow copies — the cores own opposite chips at any instant, so the two immediate writes sanitise both chips within the first frame after reset. Each writes BOOTING again after its ZP clear, restoring the wiped mirror and re-asserting on whichever chip it owns by then — covering a reset landing mid-frame where the shadow copy may straddle a swap.
CPU1's side of the same handshake is in MAD65_CPU_OS.md.
Constraints that are easy to break by accident. The reasoning behind each is in the
comments in roms/gpu_os.s.
Row-start tables must stay page-aligned. addr = $8000 + Y × 50 + X/8 needs a
multiply the 65C02 does not have, so the row start comes from two 256-byte ROM tables
baked at assemble time (row_lo[Y], row_hi[Y]). An indexed lda abs,x costs 4 cycles
only while the load stays inside one page — row_lo at $F400 and row_hi at $F500
guarantee that for every row. Moving either table off a page boundary silently costs a
cycle on every address computation in the system.
Rows 256–299 need the high-byte correction. The tables have 256 entries indexed by
the low byte of PX_Y. When PX_Y+1 = 1, add 50 to the high byte
(256 × 50 = $3200). This is why a 300-entry table was not used — 16-bit indexing does
not exist on the 65C02.
Line direction is dispatched once, not tested per pixel. line_core picks one of
four octant-specialised loops (lc_xm_R/L, lc_ym_R/L) at entry; the loops contain no
direction test. The Y direction is folded into LN_ROWSTEP, a signed 16-bit stride
(±50 for line_core, ±100 for dotline_core) added branchlessly. The step counter
lives in the X register, not a ZP variable. LN_ROWSTEP is shared between
line_core and dotline_core — safe only because the dispatch loop is single-threaded
and the two never run concurrently.
Inner-loop labels must be non-local. ca65 resets the @-local label scope at every
.local inside a macro expansion, and LC_ADVX_R/LC_ADVX_L use .local internally.
Loop-back targets therefore use function-name-prefixed labels (lc_xm_R_loop:), not
@loop:.
The sprite blitter is eight specialised paths. spr_blit_{8,16,32,64}_{ov,noov},
generated from a single SPR_BLITTER macro; the wbytes ≥ 8 paths use jmp for the
two intra-row branches the unrolled 64 px body outgrows. Picking the path once — on
width and the overlay flag — keeps every per-row branch out of the inner loop, and
the no-overlay paths compile the overlay load/shift/AND away entirely. Do not collapse
them into a runtime test.
spr_setup order matters: coarse 16-bit reject per axis first (off-right when
pixel_X ≥ 400, off-left when pixel_X + 8·W ≤ 0, mirrored for Y), so pixel_X is
inside −(8·W)…399 before any byte arithmetic and a far coordinate can never alias back
on-screen. Vertical clip then reuses the same row_lo/row_hi tables and adds +50 per
row; horizontal clip reduces to "skip any span byte outside column [0,49]", with no
partial-byte masking, because out-of-boundary bits are inert (see Two planes).
Shift direction is chosen to minimise passes and place the spill byte correctly:
N (pixel_X & 7) |
Direction | Passes | Spill byte |
|---|---|---|---|
0 |
none | 0 | right (zero) |
1–4 |
right (LSR/ROR, low→high) |
N |
right |
5–7 |
left (ASL/ROL, high→low) |
8 − N |
left |
so a non-aligned sprite never runs more than 4 shift passes. The overlay buffer stays
raw (1 = black) through the shift — shifted-in zeros must remain "no change" — and the
NOT is folded into the composite's AND step. A byte-aligned fast path was considered
and not added: sprites move to any pixel, so N = 0 is rare.
ROWCNT (≤32) comment in spr_setup is stale — sprite heights grew to 255.
Harmless today, but any change to the row loop must reason with 255, not 32.