MAD-65
Contents

MAD-65 GPU OS — Reference#

The GPU OS is the firmware in CPU2's 16 kB ROM. It boots the video subsystem, then runs one interrupt-driven render pass per VSYNC, executing the instruction list CPU1 left in ping-pong RAM. Source: roms/gpu_os.s.

Status: pre-1.0. The ROM reports V0.9 on its diagnostic screen. Several features are still outstanding and nothing has run on real hardware yet — only in madsim and the Verilator sim. Call this v0.8→v0.9: the instruction set and the PPRAM protocol are stable enough to write cartridges against, but they are not frozen until a board boots.

Comb rendering and vertical (TATE) text below are both implemented and shipped. What is still designed but not built — DOT_CIRCLE16 and the TATE rotate-90 asset script — is in MAD65_future_features.md.


Memory Map (GPU perspective)#

$0000–$77FF   RAM            scratch, stack, tables, graphics pool
$7800–$7FFF   PPRAM          ping-pong RAM 2 kB — inter-CPU communication slot
$8000–$BFFF   VRAM-image     active write buffer, hardware-switched each VSYNC
$C000–$FFFF   ROM (read) / VRAM-background (write)

VRAM note: only one VRAM buffer is accessible to the GPU at any time. Hardware switches buffers each VSYNC automatically. The GPU never knows which physical buffer it currently holds — this is the root of the two-frame background rule below.


GPU RAM Layout ($0000–$77FF, 30 kB)#

$0000–$00FF   Zero page               256 B   see below
$0100–$01FF   Stack                   256 B   65C02 hardware stack
$0200–$02FF   Scratch                 256 B   per-opcode working buffers
$0300–$06FF   Sprite definition table 1024 B  4 pages × 256 entries
$0700–$0EFF   Tile bank              2048 B   256 user tiles × 8 B
$0F00–$0FFF   (reserved)              256 B
$1000–$77FF   Graphics pool         26624 B   26 kB — sprite/tile bitmaps, uploaded via LOAD

The $0200 scratch page is shared and overlapping by design — only one opcode handler runs at a time, so the sprite row spans (SPR_BMP $0200, SPR_OVL $0210), draw_text_core's shift buffer and op_tile's glyph-pointer tables all reuse it.

Zero Page Layout ($0000–$00FF)#

$00       GPU_STATUS        ZP mirror of the PPRAM status byte
$01       NO_CPU_COUNT      consecutive frames without CPU_READY (diag mode trips at $40)
$02       VIDEO_REG_SHADOW  software shadow of the write-only VIDEO_REG
$03–$04   PPWP              instruction-list cursor (PPRAM, or DIAG_LIST in diag mode)
$05–$06   PTR               universal 16-bit VRAM / RAM pointer
$07–$0A   PX_X, PX_Y        full-res pixel coordinates, 16-bit each
$0B–$0C   PX_TMP            address-arithmetic scratch
$0D–$0E   GLYPH_PTR         current glyph row in the font
$0F       DRAW_BASE_HI      VRAM target high byte: $80 = image, $C0 = background
$10–$14   TEXT_*            column, row, scroll, visible length, glyph-row counter
$15–$21   LN_*              line state: X1,Y1,X2,Y2, DX,DY, SX,SY, ERR, CNT, MASK,
                            ROWSTEP (signed 16-bit, shared by line_core + dotline_core)
$22–$33   SPR_*             sprite blitter: width, rows, overlay flag, bit offset, shift
                            count + direction, span, column, clip flag, row counter,
                            source ptr, VRAM ptr, id, top-skip, stride, composite scratch
$34–$49   CIR_*             DOT_CIRCLE: centre, radius, midpoint walker, decision
                            variable, per-group column/row values, masks and validity
$4A       BOOT_BLINDED      one-shot — boot's BLINDER is still up; drop it on frame 1
$4B–$4E   SPR_COMB, SPR_D, SPR_VSTEP   comb-sprite ($51) state
$4F–$52   SPR_CT0/CT1       comb-sprite 16-bit scratch
$53–$64   L16_*             LINE16 / chain state: 16-bit endpoints, deltas, error,
                            step directions, the chain flag and the segment counter
$65       FRAME_BUSY        set while a command list is being dispatched — the
                            frame-overrun detector (see *Per-frame ISR*)
$66       OVERRUN_CNT       overrun frames since boot, saturating at 255
$67–$A7   PG_*              polygon ($4C–$4E): centre, angle, scale, the rotate-and-scale
                            matrix and its signs, the vertex cursor and outcodes, the
                            bisection clipper's working points, and the multiply operands
$A8–$C3   CI_*              CIRCLE16 ($4F): signed-16 full-res centre, radius, midpoint
                            walker and decision variable, the clip flag, and the per-group
                            column/row values, masks, validities and row base
$C4–$FF   free              general ROM working variables

VIDEO_REG_SHADOW is the only record of the current VIDEO_REG state (the hardware register is write-only). All code must read the shadow and write to both — never write VIDEO_REG without updating the shadow.


Ping-Pong RAM (PPRAM) Structure#

The first byte at $7800 is the status code byte (PP_STATUS); each side reads it to learn the other's state. The GPU does not process PPRAM content until the CPU has set CPU_READY. Bytes $7801 onward form the instruction list — a fixed opcode followed by zero or more argument bytes, terminated by WAI ($00).

Code Symbol Direction Meaning
$A0 CPU_BOOTING CPU → GPU CPU is initialising
$A1 CPU_WORKING CPU → GPU CPU is processing a frame
$A2 CPU_READY CPU → GPU CPU has finished writing
$B0 GPU_BOOTING GPU → CPU GPU is initialising
$B1 GPU_WORKING GPU → CPU GPU is rendering
$B2 GPU_READY GPU → CPU GPU has finished rendering

The CPU does not zero PPRAM per frame: it overwrites the list in place from $7801 and re-terminates it with WAI; stale bytes past the WAI are never read.


GPU Instruction Set#

Opcode Mnemonic Arguments (bytes following opcode) Description
$00 WAI — End of instruction list
$08 GPU_LED pattern Write bit pattern to LED_GPU_REG
$10 COPY_DIS_ON — Set COPY_DIS in VIDEO_REG
$11 COPY_DIS_OFF — Clear COPY_DIS
$12 BG_REG_ON — Set BG_REG
$13 BG_REG_OFF — Clear BG_REG
$14 BLINDER_ON — Set BLINDER (blank screen)
$15 BLINDER_OFF — Clear BLINDER
$20 CLEAR_BG — Clear VRAM-background $C000–$FFFF. Two-frame rule applies
$30 LOAD page_MSB, then 256 data bytes Copy 256 bytes into one page (page_MSB << 8). Valid: $02–$77 (RAM), $80–$BF (VRAM-image), $C0–$FF (VRAM-bg, two-frame rule). Blocked: $00, $01, $78 — data consumed but skipped
$31 RECT_BG XB, Y_LSB, Y_MSB, WB, H, GAP, then WB*H data bytes Byte-aligned rectangle onto VRAM-background. Source row r lands on row Y + r·(GAP+1) — the comb step, GAP=0 = solid. XB/WB are BYTE columns (0–49 / 1–50). Rows past 299 are sunk, not clipped: their payload is still consumed. Two-frame rule applies
$32 RECT_BG_RLE same header, then an RLE stream The same, payload compressed. See The transport block below
$33 LOAD_RLE page, LEN_LSB, LEN_MSB, then an RLE stream $30 with a compressed payload and a 16-bit length, so one command can span several destination pages. Same three blocked pages as $30, and a blocked page still consumes the stream
$40 PIXEL X_LSB, X_MSB, Y_LSB, Y_MSB Pixel on VRAM-image, full-res
$41 (removed) — PIXEL_BG — retired: a per-bit set needs read-modify-write, and VRAM-background is write-only (reads return ROM)
$42 LINE X1, Y1, X2, Y2 Solid line on VRAM-image, half-res coords
$43 LINE16 X1_LSB, X1_MSB, Y1_LSB, Y1_MSB, X2_LSB, X2_MSB, Y2_LSB, Y2_MSB Solid line on VRAM-image, full-res endpoints (0–399 / 0–299). Was LINE_BG through v0.8; the number was reused
$44 DOT_LINE X1, Y1, X2, Y2 Dotted line, image only
$45 DOT_LINES N, x0,y0, … , xN,yN Connected chain of N dotted segments (polyline), image only
$46 DOT_PIXEL X, Y One pixel at half-res coords, image only
$47 DOT_PIXELS N, x0,y0, … , xN-1,yN-1 N independent half-res pixels in one dispatch, image only
$48 DOT_CIRCLE CX, CY, R Dotted outline circle, half-res, image only. Centre must be on-screen
$4A LINES N, x0,y0, … , xN,yN Connected chain of N solid half-res segments — the solid twin of $45
$4B LINES16 N, x0.16,y0.16, … , xN.16,yN.16 Connected chain of N solid full-res segments
$49 HDOT_LINE XB, Y, NB Byte-aligned horizontal dotted rule, image only. Self-clipping
$4C DOT_POLYGON CX.16, CY.16, ANGLE, SCALE, N, dx0,dy0, … , dxK-1,dyK-1 Figure around a signed-16 centre, rotated by ANGLE and scaled by SCALE on the GPU, always clipped. N bit 7 = OPEN (polyline) / 0 = CLOSED. Dotted, half-res. 2K + 8 bytes, K = N & $7F
$4D POLYGON same as $4C The solid half-res twin
$4E POLYGON16 same as $4C The solid full-res twin: centre and offsets are read on the 400 × 300 grid
$4F CIRCLE16 CX_LSB, CX_MSB, CY_LSB, CY_MSB, R Solid circle outline around a signed-16 full-res centre that may be off screen, always clipped. R in full-res pixels, 0–255. No ANGLE/SCALE — see below
$50 SPRITE SPR_ID, X_LSB, X_MSB, Y_LSB, Y_MSB Sprite on VRAM-image, signed-16 top-left pixel, clipped on all four edges
$51 COMB_SPRITE SPR_ID, X_LSB, X_MSB, Y_LSB, Y_MSB, GAP (0–15) Comb ("ghost") sprite — same sprite, same slot, but source row r lands on screen row Y + r·(GAP+1); the GAP rows between are never written, so the background shows through. Bitmap-only: the overlay plane is ignored. GAP=0 = a plain sprite without its overlay
$60 TEXT X (0–49), Y (0–36), scroll (0–7), …string…, $00 Text on VRAM-image. Last character is never drawn — it is the scroll guard
$61 TEXT_BG same as $60 Text on VRAM-background. Two-frame rule applies
$62 VTEXT X (0–36), Y (0–49), scroll (0–7), …string…, $00 Vertical (TATE) text on VRAM-image — for a monitor stood on its side. Same (X, Y) order as TEXT; the 37 × 50 grid is its transpose. Every character is drawn, and one running past cell 36 is clipped per glyph row on the two guard cells
$63 VTEXT_BG same as $62 Vertical text on VRAM-background. Two-frame rule applies
$70 TILE X (0–49), Y (0–36), scroll (0–7), …tile ids…, $00 Tile string on VRAM-image
$71 TILE_BG same as $70 Tile string on VRAM-background. Two-frame rule applies

The CPU-side builders map 1:1 onto these opcodes — see MAD65_CPU_OS.md.

Line family — the 3 × 4 matrix#

Three line families, each offering the same four calls. Pick the family for the look and precision; the call within it for how coordinates behave at the screen edge.

single, clamped single, clipped chain (polyline) figure/polyline (rotated+scaled, GPU-side)
dotted, half-res gpu_dotline $FF1B → $44 gpu_dotline_clip $FF93 → $44 gpu_dotlines $FF1E → $45 gpu_dotpolygon $FFD2 → $4C
solid, half-res gpu_line $FF15 → $42 gpu_line_clip $FFC3 → $42 gpu_lines $FFC6 → $4A gpu_polygon $FFD5 → $4D
solid, full-res gpu_line16 $FFC0 → $43 gpu_line16_clip $FFBD → $43 gpu_lines16 $FFC9 → $4B gpu_polygon16 $FFD8 → $4E

Line family — which to use#

All line opcodes take half-resolution coordinates (0–199 × 0–149, one byte each), doubled internally to the 400×300 framebuffer — halves PPRAM usage, keeps Bresenham in 8-bit arithmetic. DOT_LINE/DOT_LINES skip every other row so half-res stepping is invisible there.

LINE ($42) draws at FULL resolution — only the endpoints are half-res (doubled once); Bresenham then steps one full-res pixel at a time (8-connected, 1 px thick, 2·max(DX,DY)+1 pixels). LINE16 ($43) is the same algorithm with signed-16 endpoints instead of doubled half-res ones, so endpoints land on any pixel rather than only even ones — the win is motion: with $42 a slowly rotating/drifting line snaps in 2 px jumps (13 distinct lines out of a 25-frame 10° sweep vs. $43's 24). Cost: 8 PPRAM bytes instead of 4, and +5–8% cycles up to 255 px on the major axis (a 16-bit Bresenham error term beyond that, +32–37% — op_line16 picks between two octant-loop sets off the major delta's high byte). Use $42 for anything static, $43 when endpoints move.

Feeding $43/$4B needs the full-res projection — this is a real trap. The CPU vector library's ordinary project/mesh_xform (and *_raw) produce half-res coordinates; doubling one to get a full-res value can only ever yield an EVEN coordinate — exactly $42's quantisation, so a $43 fed that way costs 8 extra PPRAM bytes to draw precisely what $42 draws. Project through mesh_xform_raw16 ($FFCC) or project_raw16 ($FFCF) instead — same perspective, but the quotient lands on the 400 × 300 grid and comes out odd about half the time. Cost and caveats: MAD65_CPU_OS.md.

Opcode Rendering Speed Best for
LINE ($42) Solid — full-res Bresenham, 8-connected, 1 px thick; H/V runs byte-batched baseline; H/V lines ~6–8× faster via byte fill UI frames, borders, axes, bars — anything that must look continuous
LINE16 ($43) The same line, but the endpoints are full-res (16-bit) +5–8 % over LINE up to 255 px on the major axis, +32–37 % past it Vectors whose ends move: rotating wireframes, needles, sweeps — anything that snaps visibly at 2 px. Project it with mesh_xform_raw16, or it is just $42 with more bytes
DOT_LINE ($44) Dotted — one dot per half-res step, on even coords ~1.8× faster than LINE One-off diagonal edges
DOT_LINES ($45) Dotted chain, shares vertices ~1.4–1.5× faster than the same edges sent separately, on top of DOT_LINE's gain Wireframes and any connected path
DOT_CIRCLE ($48) Dotted outline, midpoint algorithm, 8-way symmetry ~0.6 ms at R≈50, ~0.25 ms at R≈25 Reticles, radar rings, gauges, wireframe nodes
HDOT_LINE ($49) Dotted horizontal rule — constant $AA byte fill ~10–14× faster than the same span as a horizontal DOT_LINE (~13 cyc/byte vs ~194) HUD separators, grid/ruler lines, table rules, dotted underlines

All of them are image-only — there is no _BG variant of any line or pixel op. Setting an individual bit requires a read-modify-write and the VRAM-background window is write-only. Background layers are built from whole-byte writes: LOAD, TEXT_BG, TILE_BG, CLEAR_BG.

DOT_LINES shares each interior vertex between adjacent segments, so the expensive start-address computation runs once per chain; branching models are sent as several chains plus the occasional lone DOT_LINE.

DOT_CIRCLE restrictions#

HDOT_LINE encoding and self-clipping#

Three guarantees collapse the whole line to a constant memory fill: the row never changes; every dot lands on an even full-res column so a full byte of dots is always $AA; and both start and length are byte-aligned, so no mask is needed. The renderer points at the start byte and stores $AA to NB consecutive VRAM bytes.

Arg Meaning Range
XB start byte column (full-res col = XB·8) 0–49
Y half-res row — drawn on full-res row 2·Y 0–149
NB run length in bytes (1 byte = 8 full-res columns) 1–50

It is one of only two ops that validate their own arguments: XB ≥ 50 or Y ≥ 150 drops the command, and NB is clamped to 50 − XB (a resulting NB = 0 drops it). All three argument bytes are consumed before any drop, so the stream stays aligned. That bounds the highest byte it can touch at $BA65 — safely inside the image and nowhere near the $BFE0 registers. Writes are a plain store, not read-modify-write, so the odd sub-columns of each touched byte are cleared.

The CPU builder gpu_hdotline takes ordinary half-res pixel coordinates and auto-aligns them to this encoding.


The polygon family — DOT_POLYGON ($4C), POLYGON ($4D), POLYGON16 ($4E)#

A closed figure around a defined centre, rotated and scaled by the GPU, and always clipped. One command draws a whole outline:

$4C | $4D | $4E   CX.16, CY.16, ANGLE, SCALE, N, dx0,dy0, … , dxN-1,dyN-1
Field Meaning
CX, CY The centre, signed 16-bit, and it may lie off screen. Half-res units (the 200 × 150 grid extended to signed 16 — the same encoding gpu_dotline_clip already takes) for $4C and $4D; full-res (400 × 300) for $4E.
ANGLE Rotation about that centre, one byte of brad: 256 brad = a full turn, 128 = half. Same convention and the same direction as API_SIN/API_COS ($FF7B/$FF7E).
SCALE Unsigned Q0.7: 128 = 1:1, 64 = half. Shrinking only — the GPU clamps anything above 128 down to it, and treats 0 as "draw nothing".
N Bit 7 is the CLOSED/OPEN flag; bits 0–6 are the vertex count K (0–127). Bit 7 = 0 (default): CLOSED — the last vertex joins the first, so K vertices are K segments (today's behaviour, unchanged for every K an existing caller sends). Bit 7 = 1: OPEN — K vertices are K-1 segments, a true polyline through the same rotate+scale matrix.
dx, dy Signed byte offsets from the centre, unrotated and unscaled — the raw shape. Half-res for $4C/$4D, full-res for $4E.

CLOSED is a figure around a point; OPEN is a rotated, scaled polyline — N's top bit picks which. No _CLIP suffix, because the anchor is signed-16 and may sit off screen, so the family is always clipped either way. For a path with no rotate/scale matrix, or more than 127 vertices, use the chains ($45/$4A/$4B) instead.

It generalises DOT_CIRCLE ($48) (already anchored at its centre, but with a byte anchor that must be on screen) — $48 is untouched.

What $4E unlocks#

A CPU-side quarter-square multiply keeps a transformed vertex at half-res precision at best (doubling only ever yields even coordinates). The GPU instead interprets the same ±127 offset as full-res directly: a figure up to 254 px across (the limit — ±128 no longer fits a signed byte), each rotated offset rounding to a whole full-res pixel — full-res figures cannot be produced any other way. $4E costs roughly twice $4C's pixels (full resolution vs. every-other-row) plus LINE16's +5–8% for 16-bit endpoints.

Cost#

Measured in py65 on a 12-vertex, radius-48-half-res figure. "Transform + clip" is what moved onto the GPU versus the same shape pre-transformed and sent as a chain.

wholly on screen crossing an edge rejected transform + clip
$4C DOT_POLYGON 26 923 (11.3% of a frame) 21 271 481 8 000 (666/vertex)
$4D POLYGON 39 513 (16.6%) 26 040 483 8 310 (692/vertex)
$4E POLYGON16 41 387 (17.4%) 28 934 483 7 481 (623/vertex)

A frame is 237 404 cycles; cost scales linearly with vertex count, so level-of-detail (fewer vertices for a small figure) is entirely CPU1's to pull — there is no shape table or SHAPE_ID on the GPU. An off-screen figure costs ~480 cycles flat, about a sixtieth of drawing it. This is a transfer, not a speed-up — the same algorithm costs the same on either 65C02; what moves is which processor is idle when it runs.

The GPU folds rotation and scale into one matrix built once per figure (so the scale rides free inside the rotation's products) and clips with a Cohen–Sutherland outcode test per segment, sharing the chain's joint saving between two ends that are both inside. It clips, never clamps — a clamped corner would bend every segment touching it, so true signed-16 geometry is kept until the cut, a division-free bisection (sixteenths of a unit, ≤4 passes) snapped onto the crossed boundary.

Restrictions#

Trig, and the one-unit difference from API_SIN#

The GPU carries its own sine/quarter-square tables (1152 B at $F300/$F600–$F9FF — the two processors can't see each other's ROM). The quarter-square multiplier (a·b = f(a+b) − f(|a−b|), f(x) = ⌊x²/4⌋) costs ~60 cycles/product against ~130 for bit-serial. The sine table is scaled 0–128, not the CPU OS's 0–127 Q0.7: that makes ANGLE=0, SCALE=128 an exact identity and a quarter turn exact, at the cost of a 1-unit-in-128 disagreement with API_SIN/API_COS on the same angle — never in direction (both turn +X toward +Y, clockwise on screen).

Tests#

roms/test_polygon.py: exact pixel-set equality against the identity case ($4C==$45, $4D==$4A, $4E==$4B), an independent Liang–Barsky clip for edge-crossing figures, an N=1 probe of the rotate/scale matrix directly, and rotation direction against API_SIN/API_COS.


CIRCLE16 ($4F) — the circle with an anchor that may be off screen#

$4F   CX_LSB, CX_MSB, CY_LSB, CY_MSB, R                      (6 bytes)

The circle counterpart of the polygon family: a solid outline around a signed-16 full-res centre that is allowed to sit off screen, and therefore always clipped. R is a radius in full-res pixels, 0–255; R = 0 draws nothing. Builder: gpu_circle16 ($FFDB).

$48 DOT_CIRCLE is untouched and is not superseded — it remains the dotted, half-res, on-screen-only call, at about half the cycles ($4F runs the same midpoint walk at full resolution, which is what makes its outline solid).

No ANGLE, no SCALE, no half-res twin. ANGLE does nothing to a circle (the polygon family's rotate/scale matrix has nothing to compose here); SCALE would ride in no existing product (one extra multiply + stream byte to save the caller one it already does per object — a composite object still works from one struct, only the anchor shared); and a half-res form would cost the same 6 stream bytes and cycles as this one, buying only a radius range past 255 px in exchange for an even-only centre. The gap that would earn an opcode is the dotted twin, DOT_CIRCLE16 — see MAD65_future_features.md.

Solid and clipped exactly. 8-connected, 1 px thick, same octant walk as $48 but at full resolution ($48 doubles the half-res walk instead, hence dotted). Clipping is decided once from the bounding box:

Bounding box What happens Cost
misses the screen nothing drawn 200 cycles
wholly inside fast path, no per-point range tests see below
straddles an edge per-point tests every column/row +8–15%

Clipped per point, not geometrically — no intersection to compute — so $4F is the one clipped primitive whose output is exactly predictable: roms/test_circle16.py holds it to midpoint_walk(cx,cy,r) ∩ screen with no tolerance.

Radius (full-res px) wholly on screen half off an edge
10 6 426 (2.7% of a frame) 7 336
50 27 571 (11.6%) 31 626
150 105 039 (44.2%) 93 074

Linear in R (every point is a read-modify-write). A 100 px-radius circle is a quarter of a frame and should be a deliberate choice; $48 at the same visible size costs about half.


⚠ Coordinate validation is CPU1's responsibility#

The GPU drawing routines do not validate coordinates. Out-of-range values are not caught, clipped or rejected — the GPU computes whatever address the formula gives and writes there. CPU1 must ensure every coordinate is in range before it reaches the instruction list.

Instruction family Parameter Legal range Notes
PIXEL X, Y 0–399 / 0–299 (16-bit) Full-resolution
LINE, DOT_LINE, DOT_LINES, DOT_CIRCLE X, Y 0–199 / 0–149 Half-res, doubled internally
DOT_PIXEL, DOT_PIXELS X, Y 0–199 / 0–149 Same space as lines
TEXT, TEXT_BG, TILE, TILE_BG column, row 0–49 / 0–36 Tile grid units (50 × 37 cells)
SPRITE, COMB_SPRITE X, Y signed 16-bit Top-left pixel, (0,0) = screen origin; GPU clips all four edges
COMB_SPRITE GAP 0–15 Masked with $0F; the high nibble is reserved and ignored
DOT_CIRCLE CX, CY 0–199 / 0–149 Centre must be on-screen (early-out if violated)
HDOT_LINE XB, Y, NB 0–49 / 0–149 / 1–50 Self-clipped — see above
DOT_POLYGON, POLYGON CX, CY signed 16-bit, half-res units May be off screen — the GPU clips every edge
POLYGON16 CX, CY signed 16-bit, full-res units May be off screen — the GPU clips every edge
$4C–$4E dx, dy −127…+127 (signed bytes) −128 is clamped to 127; the figure is at most 254 units across
$4C–$4E SCALE 0–128 (Q0.7, 128 = 1:1) Shrinking only; above 128 clamped to 128, 0 draws nothing
CIRCLE16 CX, CY signed 16-bit, full-res units May be off screen — the GPU clips every point
CIRCLE16 R 0–255, full-res pixels R = 0 draws nothing

What happens with illegal values#

The address formula for pixel and line instructions is:

VRAM address = $8000 + Y_full × 50 + X_full / 8

VRAM-image ends at $BFDF. Immediately above it sit write-only hardware registers: LED_GPU_REG at $BFC0 and VIDEO_REG at $BFE0 (BG_REG bit 0, BLINDER bit 1, COPY_DIS bit 2, SHADOW_MODE bit 3). If Y_full ≥ 326 the computed address reaches $BFC0 and a pixel write latches random bits into one of them — BLINDER blanks the screen instantly, BG_REG switches the displayed layer, COPY_DIS kills the background copy, and clearing SHADOW_MODE un-maps the shadow RAM the GPU is executing from. All persist until the next valid VIDEO_REG write, and the last one does not survive to make one. For half-res instructions the threshold is Y ≥ 163. Values above the legal maximum but below the threshold write into the unused VRAM tail — invisible, but still wrong.

The two register ranges are decoded separately (VIDEO_REG needs A7 & A6 & A5), so GPU_LED is safe to issue at any time, including during rendering. See MAD65_address_decoder.md §3/§4.

Why the GPU does not validate#

Bounds checking costs cycles on every draw call for coordinates that are almost always valid; validation belongs in CPU1's scene-composition logic, where it runs once a frame instead of once a pixel. DOT_CIRCLE's centre check, SPRITE's clipping, HDOT_LINE's row-end clamp and the polygon family's clip are exceptions only because those routines need the range information to drive their own inner loops (bounding box, clip region, byte count) — a per-segment clip in particular runs once per SEGMENT, not once per pixel, so it costs the same either processor and only changes which one is idle. Pure drawing instructions like PIXEL/LINE have no such internal need and stay unchecked.


The transport block — RECT_BG ($31), RECT_BG_RLE ($32), LOAD_RLE ($33)#

Three opcodes added together because they are one job. None of them is drawing: they move bytes from the PPRAM command stream into GPU memory, which is why they live beside LOAD in the $30 block and why they are dispatched from dispatch_rare rather than the fast chain — a background load is a handful of commands a frame across ~17 frames, not a per-frame draw. The fast chain keeps its last free slot for a future drawing opcode, where 3 cycles a command would actually matter.

What LOAD could not do#

$30 is a linear 256-byte copy to a page-aligned destination, and a page straddles 5.12 screen rows — structurally row-agnostic, unable to express an arbitrary rectangle or a comb. $31 takes a byte column and a row and walks the destination by 50·(GAP+1).

The RLE format#

ctrl $01–$7F   the next ctrl bytes are literals
ctrl $80–$FF   repeat the ONE byte that follows, (ctrl − $80) + 2 times ($80 = 2 … $FF = 129)
ctrl $00       never emitted

A repeat run is 2 stream bytes for up to 129 output bytes; a literal run costs 1 byte of overhead per 127, so the worst case on incompressible data is +1/127 ≈ +0.8 %. Measured on this repository's own 1-bit art it returns 1.4×–2.3×, which is roughly three quarters of what a general-purpose compressor finds, for about a hundred bytes of decoder.

There is no end marker and no length header. The decoder is driven by an output count the consumer already knows — $32 consumes WB·H, $33 takes its count as an argument. That keeps the decoder state to three bytes, and an over-long final run is harmless.

The producer's obligation, and it is NOT checked at runtime. A stream too short for its output count is read straight past its end and into the next command, and the dispatch loop desynchronises — arbitrary bytes then execute as opcodes. For $32 that means every band must encode exactly WB·H bytes, so a picture whose height is not a multiple of H needs its last band padded. roms/rle.py does this. Checking it here would cost a byte budget in every command to catch a build-time mistake.

Why compression lives on this side of PPRAM#

The background layer is transport-bound, not blit-bound (~3.5 pages/frame through the ping-pong window), so compressed bytes must cross PPRAM and expand here — decompressing on CPU1 would send raw bytes through the one bottleneck that matters and win nothing. The mirror case — a payload sinking in CPU1's own RAM — never touches PPRAM, so CPU1 carries its own decoder (api_decrunch). Compress at the source, decompress at the sink.

The two decoders are deliberately different codecs. This plane is write-only, so a decoder here can never look back at what it has written and has to be stateless — RLE. CPU1's sink is ordinary RAM, so api_decrunch runs LZ (roms/lz.py), which copies from its own output and beats RLE on everything CPU1 loads (the demo's VGM song 6.3×, where RLE gets 1.00×). Until 2026-09-11 both sides ran this RLE.

The bottom edge is the one thing that is checked#

Background writes go through DRAW_BASE_HI = $C0, so the address is $C000 + row·50 + XB in a 16-bit pointer. The framebuffer's last byte is $FA97. Past $FFFF the pointer wraps to $0000 — zero page and the stack, not a harmless corner of VRAM. The unused tail $FA98–$FFFF is 1384 bytes, which is 27 rows at GAP = 0 and exactly one row at GAP = 15: the comb multiplies the row step by up to 16 and eats the margin.

Every row is therefore bounds-tested and an out-of-range row is sunk — its payload is read and discarded. Sinking rather than stopping is the point: with an RLE payload you cannot skip forward without decoding, and stopping early would leave the rest of the stream in PPRAM for the dispatch loop to execute.

The horizontal edge is deliberately not checked. XB + WB > 50 smears one row into the start of the next — cosmetic, confined to the background plane, and a loader contract rle.py checks offline.

Cost#

About 45 cycles a payload byte (the source vector, the fetch and the store). During a background load CPU1 can push ~900 payload bytes a frame, so roughly 40 k cycles — about 17 % of the frame, on frames where the game is loading rather than playing. A worked example with real numbers is in carts/rect_bg_test/.


Writing to the background — two consecutive frames#

Any instruction that writes VRAM-background must be sent twice, on two consecutive frames: CLEAR_BG ($20), TEXT_BG ($61), TILE_BG ($71), and LOAD ($30) to a $C0–$FF page.

Why: VRAM is double-buffered in hardware, and a flip-flop swaps which physical chip is "the background" on every VSYNC. The GPU never knows which one it holds, so a single write lands in only one buffer; the next frame the hardware swaps to the other, which still holds the old content — the change flickers at 30 Hz or vanishes. Writing the same data on two consecutive frames reaches both buffers, after which it is stable.

Image-layer instructions do not need this. The image buffer is rebuilt every frame, so a one-frame write is correct there. Hardware copies the background onto the image each frame automatically, so the background only needs redrawing when its content actually changes — but when it does, across two frames.

The CPU OS automates this with an argument-capture replay engine, so a game issues a background write once — see MAD65_CPU_OS.md.


Graphic Data Structures#

Tiles#

Example (letter 'A'):
  00011100 = $1C      01111111 = $7F
  00110110 = $36      01100011 = $63
  01100011 = $63      01100011 = $63
  01100011 = $63      00000000 = $00

Font (TEXT)#

Horizontal scroll & the hidden last character#

TEXT/TEXT_BG take a scroll argument (0–7). The whole rendered row buffer is treated as one wide binary number and shifted left by scroll pixels, bits carrying right-to-left between adjacent cells, so the shift is pixel-continuous across the string.

The last character of the string is never written to VRAM — it is the scroll buffer. Its high bits do carry into the rightmost visible cell during the shift, so the next character scrolls in smoothly from the right. A left-scrolling marquee is therefore trivial: hold a window of N+1 characters (N visible + 1 guard), increment scroll 0→7 one pixel per frame, and when it wraps advance the string pointer by one character. Static text must end with a space.

The visible draw count TEXT_LEN is total − 1, capped at the 50-cell row buffer and clipped at the right screen edge (50 − column). scroll = 0 skips the shift entirely (static-text fast path). A 51-byte string fills all 50 columns; longer strings clip, never overflowing onto the next row.

⚠ A scrolling string must stop one cell short of the right edge. Phase 0 stores at most 50 cells, so a 51-character string at column 0 gives 50 visible cells with the guard pushed out of the buffer — TEXT_LEN == stored, no +1, and zeros shift into column 49. The text still fills the width, but each new character pops into existence instead of sliding in, once every 8 frames. A marquee therefore wants 50 characters (49 visible + guard), or more generally visible ≤ 49 − column. Filling all 50 columns and smooth-scrolling are mutually exclusive on this path.

This is a TEXT/TEXT_BG rule only. VTEXT/VTEXT_BG have no scroll guard — their scroll is an address offset rather than a bit-shift chain and every character is drawn — so a vertical marquee may use the full 37-cell width (see carts/tate_test).

Vertical (TATE) font and text (VTEXT)#

VTEXT ($62) / VTEXT_BG ($63) draw text for a monitor turned 90° clockwise (madsim's F12 mirrors the turn on screen). They are the only orientation-aware routines: a quarter turn maps the 400×300 framebuffer onto itself, and everything else already works in framebuffer coordinates, so sprites, lines, tiles and LOADs just need pre-rotated assets, no code change. See MAD65_future_features.md.

The grid is 37 × 50 — the exact transpose of TEXT's 50 × 37, given in the same (X, Y) order and directions as every other cell-grid call:

argument range on screen in the framebuffer
X — char cell 0–36 left → right, 37 cells (296 of 300 px, 2 px margin) fb-y = 290 − 8×X (decreasing)
Y — line 0–49 top → bottom, 50 lines (400/8 exact) byte column Y (fb-x)
scroll 0–7 shifts the string LEFT, 1 px at a time — same direction as TEXT adds 50 × scroll to the address

Screen position is 2 + 8×X − scroll px from the left edge — the identical marquee recipe TEXT uses, but here it's free: along fb-y a pixel is a whole VRAM row, so scroll is an address offset, not a bit-shift chain, and VTEXT draws every character (no hidden last cell like TEXT). roms/test_vtext.py pins the shared direction. To scroll the other way, count scroll down 7→0 instead — it's an offset, not a direction (see carts/tate_test/).

Cells 37/38 exist but can't be named (X > 36 draws nothing) — reached only by a string longer than its line, they're what lets a character enter from the right edge one pixel column at a time. Two guard cells, not one, because a cell is 8 px wide and scroll spans 0–7. Clipping is per glyph row, right edge only:

Glyphs come from font_data_v at $F000, generated at build time from roms/fonts.s by roms/gen_font_v.py — never hand-edited. It sits exactly $0C00 below font_data (linker-asserted), letting draw_vtext_core reuse font_glyph_lo[] with only the high byte changed.

Builders gpu_vtext/gpu_vtext_bg ($FFB4/$FFB7) pass the cell through unclamped (clamping a signed coordinate would destroy $FE/$FF); the GPU clips it.

Sprites#

Page Address range Content
Page 0 $0300–$03FF SPR_TYPE
Page 1 $0400–$04FF SPR_PTR_LSB
Page 2 $0500–$05FF SPR_PTR_MSB
Page 3 $0600–$06FF SPR_HEIGHT

Type & height#

Width and overlay come from the type byte: low nibble = width in bytes W (1 = 8 px, 2 = 16 px, 4 = 32 px, 8 = 64 px); bit 4 ($10) set → no overlay plane, clear → overlay present.

Height is independent, in SPR_HEIGHT ($0600 + SPR_ID): any value 1–255 rows. It is not derived from width, so non-square sprites are allowed (an 8×64 column, a 32×8 bar). Height 0 marks an undefined slot and renders nothing — boot zeroes the table, so unused ids are inert. The blitter costs the same per row regardless of height, so tall sprites cost proportionally more only because they have more rows.

Type $01 $11 $02 $12 $04 $14 $08 $18
Width 8 px 8 px 16 px 16 px 32 px 32 px 64 px 64 px
Overlay yes no yes no yes no yes no

64 px (W=8) is the widest the type byte can encode: W=16 (128 px) would be $10, which collides with the overlay-absent bit. 128 px would need a re-encoded type byte.

Pixel data format#

1 bit per pixel, packed 8 px/byte, MSB = leftmost — same bit order as tiles and glyphs. Stored interleaved by row:

Two planes#

Plane Bit = 1 Bit = 0 Operation
Bitmap white transparent OR (screen \| bitmap)
Overlay black no change AND NOT (& ~overlay)

The composite is a single read-modify-write per byte:

screen = (screen OR bitmap) AND NOT(overlay)

The bitmap draws the white silhouette; the overlay punches black detail on top. Because both planes treat a 0 bit as "leave alone", a sprite shifted to a non-byte-aligned column needs no boundary masking — the zero bits shifted in on the spill side are inert, and partially-visible edge bytes come out correct without per-bit clipping.

Comb ("ghost") sprites — COMB_SPRITE ($51)#

$51 takes SPRITE's five arguments unchanged in meaning and order, plus a sixth byte GAP (0–15, masked with $0F). With D = GAP+1, source row r is blitted to screen row Y + r·D; the GAP rows in between are never written, so whatever the hardware copied under the scene stays visible. On-screen height is (ROWS−1)·D + 1, i.e. the same sprite data covers up to 16× its own height.

Budgeting it (measured on carts/comb_test, unit = one 64 px-wide slot, one source row; a W-px-wide, H-row sprite is ceil(W/64) × H of them): ~330 cycles/slot-row byte-aligned, ~570 at the worst bit offset (spr_setup picks the cheaper shift direction, so it's a tent peaking at offset 4, not a ramp). A 256×83 source filling the screen at GAP=1 runs 46–79% of a frame — a comb sprite that wide is a scene, better placed in the background layer unless it's the only thing moving; a 128×41 one at 12–20% is an ordinary moving object, and GAP then adds height for free (same 12% at 161 px tall, GAP=3).

Clipping is spr_comb_vclip, which differs from $50's vertical clip in exactly two places, both of which are where the bugs would be:

The division that both of these need (div_small, 16 ÷ 1–16) runs at most twice per sprite and never per row; the top-clip call only happens when the sprite actually hangs off the top edge. Tests: roms/test_comb_sprite.py, which checks every residue class of top_skip mod D and asserts on every case that nothing was written at or past $BA98. Working example: carts/comb_test/ — a ghost planet drifting over background captions that stay readable straight through it, with GAP on the joystick and an end-to-end harness that measures the see-through property rather than eyeballing it.


Per-frame ISR#

The GPU never returns from its interrupt — VSYNC is the frame trigger, so frame_isr resets SP on entry and ends in WAI.

VSYNC interrupt fires
  Reset SP ← $FF               (discard the hardware-pushed frame; there is no caller)
  If FRAME_BUSY:               ← the previous frame overran; see the policy below
    OVERRUN_CNT += 1           (saturating)
  FRAME_BUSY ← 0
  CLI                          (a late VSYNC must not be blocked)

  If NO_CPU_COUNT >= $40:      ← diagnostic mode check, before anything else
    PPWP ← DIAG_LIST ($FA00)
    goto dispatch_loop         (PPRAM status check skipped entirely)

  PPWP ← PPRAM ($7800)         ← normal mode
  Read status byte via PPWP    (PPWP advances to $7801)
  If status ≠ CPU_READY:
    NO_CPU_COUNT += 1
    status ← GPU_READY
    WAI                        (nothing rendered this frame — a visible blink)
  NO_CPU_COUNT ← 0
  If BOOT_BLINDED:             ← one-shot: first real frame drops boot's BLINDER
    clear BLINDER, clear BOOT_BLINDED
  status ← GPU_WORKING
  FRAME_BUSY ← $80             (armed: a VSYNC before the end of the list is an overrun)
                               (PPWP already at $7801 = first opcode)
dispatch_loop:
  Read opcode via PPWP, increment PPWP
  If opcode = $00 (WAI):
    FRAME_BUSY ← 0             (finished inside its VSYNC period)
    status ← GPU_READY
    WAI
  Dispatch to the opcode's routine
    Routine reads its own arguments via PPWP and advances it
  goto dispatch_loop

Convention: whenever a byte is read via PPWP, increment PPWP immediately, so it always points at the next unread byte. Opcode routines end with JMP dispatch_loop, not RTS — this keeps the call depth flat.

Dispatch is a linear CMP chain ordered by call frequency, hottest first (dotted geometry → pixel/sprite → text/tile → solid line → polygon → circle → LOAD → WAI → rare control ops). Cost is 30 + N×4 cycles for the Nth entry. All BEQ targets are 3-byte JMP trampolines listed in the same order as the comparisons, because BEQ only reaches ±128 bytes; the assembler errors if one falls out of range.

The chain is capped at 29 entries. The first BEQ has to reach the first trampoline, which is 117 bytes ahead at 29 entries and out of range at 30. The polygon family took the last three slots and CIRCLE16 a fourth, so COPY_DIS_ON/_OFF, BG_REG_ON/_OFF and GPU_LED moved into dispatch_rare — a BNE+JMP tail with no reach constraint, reached by falling off the end of the chain. All five are written at most once per state change, never inside a per-frame loop, so the extra JMP costs nothing that matters. The chain currently holds 28. A new opcode goes into dispatch_rare, or earns its way up by displacing something.

Frame-overrun policy#

FRAME_BUSY is set just before the dispatch loop is entered and cleared when the list's WAI is reached. Finding it set on entry to frame_isr means a VSYNC arrived while the previous frame was still drawing: that frame overran its budget.

The policy is "the new frame wins" — there is no choice about it. The ping-pong hardware has just swapped the GPU's 2 kB window, so the abandoned dispatch's command list is gone; resuming would splice one list into the middle of another (HUD text executed as opcodes: ' ' is CLEAR_BG, '0' is LOAD, and further on lie VIDEO_REG and BLINDER). The SP reset at the top of frame_isr has already discarded the old frame, and it never returns.

A viewer sees one partially drawn image, for exactly one frame — the hardware background copy re-lays every frame, so a truncated image never persists: consistent overrun produces torn frames, not accumulating wreckage, and recovers the moment load drops. CPU1 is not told (no handshake on GPU_READY); the GPU only counts: OVERRUN_CNT saturates at 255, read by nobody but a bench or debugger. "Finish the frame you started, drop the one you missed" is not available — it would need the in-flight list to survive the VSYNC, and the swap is unconditional hardware.


GPU ROM Layout ($C000–$FFFF, 16 kB)#

$C000–$C002   JMP stub — RESET → boot_main
$C003–$C005   JMP stub — IRQ   → frame_isr
$C006–$????   Boot procedure                            (boot_main)
$????–$EFFF   Frame handler, dispatch loop, op_* routines, helpers
$F000–$F2FF   Rotated font (TATE)           768 B       96 glyphs × 8 B (FONT_V segment)
$F300–$F37F   Polygon |sin| table           128 B       pg_sinmag, 0-128 scale, brad.
                                                        128 entries, so 128-alignment is
                                                        all an indexed load needs
$F380–$F3DD   Vertical text tables           94 B       vtext_col_lo/hi + vtext_scr_lo/hi
$F3DE–$F3FF   Reserve                        34 B       filled $EA
$F400–$F5FF   VRAM row-start tables         2×256 B     row_lo ($F400) / row_hi ($F500)
$F600–$F9FF   Polygon quarter-square mult   4×256 B     pg_qs_lo/hi ($F600/$F700),
                                                        pg_qr_lo/hi ($F800/$F900) = the
                                                        same table + 64 (rounding term)
$FA00–$FAFF   Diagnostic instruction list   256 B       82 B used (gpu_diag_list.s)
$FB00–$FBBF   Font glyph pointer tables     2×96 B      FONT_GLYPH_LO ($FB00) / _HI ($FB60)
$FBC0–$FBFF   Reserve                        64 B       filled $EA
$FC00–$FEFF   Font definitions              768 B       96 glyphs ($20–$7F) × 8 B
$FF00–$FF49   Text row-start tables         2×37 B      TEXT_ROW_LO ($FF00) / _HI ($FF25)
$FF4A–$FFB1   Default sprite 0              104 B       32×26, no overlay (SPRITE0 segment)
$FFB2–$FFF9   Reserve                        72 B       filled $EA
$FFFA–$FFFB   NMI vector                                (unused — points to irq_stub)
$FFFC–$FFFD   RESET vector                              → $C000
$FFFE–$FFFF   IRQ/BRK vector                            → $C003

Segment starts are pinned in roms/gpu.cfg; the CODE end moves with every build (see gpu_os.map). CODE owns one contiguous run, $C000–$EFFF, and currently ends at $E9AA — 1621 bytes free. Data segments are page-aligned only where an indexed load can actually cross a page (row_lo/row_hi, the four quarter-square pages); anything reached through a pointer or bounded inside a page is packed tight instead. Four small islands remain that CODE cannot reach — 344 B total, named as reserve in gpu.cfg rather than left looking like an oversight.

Two placements are load-bearing, and a linker assert fails the build if either is disturbed:

The two JMP stubs give the hardware vectors fixed targets while the real handlers live anywhere in the image. row_lo/row_hi, the quarter-square pages and the font are page-aligned deliberately — see Implementation invariants.

Diagnostic mode#

If frame_isr counts 64 consecutive frames without CPU_READY (~1 s), NO_CPU_COUNT reaches $40 and the GPU enters diagnostic mode. From then on frame_isr points PPWP at DIAG_LIST ($FA00) instead of PPRAM and jumps straight into the dispatch loop, which is oblivious to where its opcodes come from. NO_CPU_COUNT is never reset in diag mode — the CPU is not expected to recover; a missing CPU signal after boot is fatal.

The list (roms/gpu_diag_list.s) uses the same opcode format as PPRAM but has no status byte — opcodes begin at byte 0 — and must fit 512 B and end with WAI. Current content: BLINDER_OFF (the diagnostic screen reveals itself, since boot leaves the screen blinded), sprite 0, and three TEXT lines — MAD-65 PROTOTYPE V0.9 / CPU NOT READY / GPU IN DIAGNOSTIC MODE. It doubles as a graphics self-test with no CPU involvement; widening it (border, full ASCII sweep, VIDEO_REG flag cycling) is a matter of adding opcodes to that file.


GPU Boot Procedure#

SEI + CLD                          disable interrupts, force binary mode
VIDEO_REG ← BLINDER set, COPY_DIS + BG_REG clear    (still SHADOW_MODE=0 boot mode)
LED_GPU_REG ← $01                  stage 0: boot started
PPRAM[$7800] ← GPU_BOOTING ($B0)   + ZP mirror — the FIRST PPRAM write of boot,
                                   deliberately BEFORE the ~1-frame shadow copy
NO_CPU_COUNT ← 0                   (the ZP clear below re-zeroes it anyway)
init stack (S ← $FF)
copy ROM $C000–$FFFF → shadow code-RAM    boot mode: read hits ROM, write hits shadow.
                                   SHADOW_MODE stays 0 until just before the first
                                   $C000–$FFFF write (the VRAM-background clear)
init zero page                     wipes GPU_STATUS / VIDEO_REG_SHADOW — intentional
PPRAM[$7800] ← GPU_BOOTING ($B0)   + ZP mirror, AGAIN: re-establishes the wiped mirror
                                   and re-asserts the byte on the chip owned by now
init RAM ($0100–$77FF)
LED_GPU_REG ← $03                  stage 1: RAM cleared
install default sprite 0 into the definition table
VIDEO_REG ← BLINDER | SHADOW_MODE  enter run mode: code now runs from shadow RAM and
                                   $C000–$FFFF writes reach VRAM-background. Must land
                                   AFTER the ZP clear (which wiped VIDEO_REG_SHADOW)
                                   and BEFORE the first $C000+ write below
clear VRAM-background × BG_CLEAR_PASSES   back-to-back loop, NO waits between
re-assert BLINDER                  (local invariant: blind while stamping the bg)
draw build-date stamp              → the background bank owned right now
wait one VSYNC                     bg bank ownership swaps
draw build-date stamp              → the other bank
wait one VSYNC                     let the bg→image copy populate both image banks
LED_GPU_REG ← $07                  stage 2: VRAM cleared and stamped
BOOT_BLINDED ← 1                   arm the reveal — the screen stays DARK for now
PPRAM[$7800] ← GPU_READY ($B2)     + ZP mirror
LED_GPU_REG ← $0F                  stage 3: GPU ready
CLI                                enable interrupts
WAI

The screen is not unblanked at the end of boot. Revealing here would flash the bare datestamp on an empty background for every frame the CPU takes to come up. Instead boot arms BOOT_BLINDED, and the reveal happens on the first CPU_READY frame, inside frame_isr, just before the first PPRAM list is drawn. It is a one-shot, so a game that later drives BLINDER itself is never overridden. The diagnostic screen reveals itself via DIAG_LIST's own BLINDER_OFF.

Why the background clear is a tight loop. Only the background is cleared — the video circuit copies background → image every frame even while BLINDER is set, so a clean background blanks the display on its own. The clear must fully zero both double-buffered halves; one pass takes ~0.6 of a frame, so must not be phase-locked to VSYNC ("clear; wait exactly one frame; repeat" made every pass start at the same phase, leaving one buffer with a permanently orphaned region — a 30 Hz flicker after reset, since SRAM retains pre-reset content). The fix is a back-to-back loop with no waits: ownership still flips once per VSYNC, but the clears land at every sub-frame phase, so each buffer gets at least one complete pass. BG_CLEAR_PASSES is 12 (~0.12 s); measured, both buffers are zero by frame ~2–3 — wide margin.

The build-date stamp is drawn after the clear loop (so no pass can erase it) and before the reveal. It lives in the background, so the free hardware copy re-paints it every frame and it survives into whatever the CPU draws, as long as the program leaves the top-right corner alone. It is tiny (14 cells × 8 rows) and finishes well inside one bank's ownership window, so unlike the full clear it never straddles a swap — hence the simple draw / wait one VSYNC / draw again. BUILD_DATE is a "YYMMDD" literal the Makefile regenerates into build_date.inc before every assembly.

Sanitising the status handshake (why GPU_BOOTING is written twice)#

PPRAM survives a warm reset, so after a reset both ping-pong chips still hold the previous run's status bytes — one ends with CPU_READY, the other with GPU_READY. Each core can only write the chip it currently owns (they swap every VSYNC), so a single BOOTING write sanitises only one of the two. A surviving stale byte false-triggers either direction: stale GPU_READY lets the CPU's boot-tail wait pass while the GPU is still clearing RAM (losing the CPU's first command lists, including one-shot LOADs); stale CPU_READY makes the GPU's first frame_isr execute a leftover pre-reset list.

The defence: both cores announce BOOTING immediately at boot entry, before their slow ~1-frame shadow copies — the cores own opposite chips at any instant, so the two immediate writes sanitise both chips within the first frame after reset. Each writes BOOTING again after its ZP clear, restoring the wiped mirror and re-asserting on whichever chip it owns by then — covering a reset landing mid-frame where the shadow copy may straddle a swap.

CPU1's side of the same handshake is in MAD65_CPU_OS.md.


Implementation invariants (drawing routines)#

Constraints that are easy to break by accident. The reasoning behind each is in the comments in roms/gpu_os.s.

Row-start tables must stay page-aligned. addr = $8000 + Y × 50 + X/8 needs a multiply the 65C02 does not have, so the row start comes from two 256-byte ROM tables baked at assemble time (row_lo[Y], row_hi[Y]). An indexed lda abs,x costs 4 cycles only while the load stays inside one page — row_lo at $F400 and row_hi at $F500 guarantee that for every row. Moving either table off a page boundary silently costs a cycle on every address computation in the system.

Rows 256–299 need the high-byte correction. The tables have 256 entries indexed by the low byte of PX_Y. When PX_Y+1 = 1, add 50 to the high byte (256 × 50 = $3200). This is why a 300-entry table was not used — 16-bit indexing does not exist on the 65C02.

Line direction is dispatched once, not tested per pixel. line_core picks one of four octant-specialised loops (lc_xm_R/L, lc_ym_R/L) at entry; the loops contain no direction test. The Y direction is folded into LN_ROWSTEP, a signed 16-bit stride (±50 for line_core, ±100 for dotline_core) added branchlessly. The step counter lives in the X register, not a ZP variable. LN_ROWSTEP is shared between line_core and dotline_core — safe only because the dispatch loop is single-threaded and the two never run concurrently.

Inner-loop labels must be non-local. ca65 resets the @-local label scope at every .local inside a macro expansion, and LC_ADVX_R/LC_ADVX_L use .local internally. Loop-back targets therefore use function-name-prefixed labels (lc_xm_R_loop:), not @loop:.

The sprite blitter is eight specialised paths. spr_blit_{8,16,32,64}_{ov,noov}, generated from a single SPR_BLITTER macro; the wbytes ≥ 8 paths use jmp for the two intra-row branches the unrolled 64 px body outgrows. Picking the path once — on width and the overlay flag — keeps every per-row branch out of the inner loop, and the no-overlay paths compile the overlay load/shift/AND away entirely. Do not collapse them into a runtime test.

spr_setup order matters: coarse 16-bit reject per axis first (off-right when pixel_X ≥ 400, off-left when pixel_X + 8·W ≤ 0, mirrored for Y), so pixel_X is inside −(8·W)…399 before any byte arithmetic and a far coordinate can never alias back on-screen. Vertical clip then reuses the same row_lo/row_hi tables and adds +50 per row; horizontal clip reduces to "skip any span byte outside column [0,49]", with no partial-byte masking, because out-of-boundary bits are inert (see Two planes).

Shift direction is chosen to minimise passes and place the spill byte correctly:

N (pixel_X & 7) Direction Passes Spill byte
0 none 0 right (zero)
1–4 right (LSR/ROR, low→high) N right
5–7 left (ASL/ROL, high→low) 8 − N left

so a non-aligned sprite never runs more than 4 shift passes. The overlay buffer stays raw (1 = black) through the shift — shifted-in zeros must remain "no change" — and the NOT is folded into the composite's AND step. A byte-aligned fast path was considered and not added: sprites move to any pixel, so N = 0 is rare.


Open items#