Per-vertex GS fog end-to-end (gs_stub emit incl. persp_emit5, gs_prim_list_feeder XYZ2->XYZF2 on PRIM.FGE, gs_make_sh3_scheduler_fixture.py F/FGE packing), new fog TBs, fidelity attribution tooling. Functional baseline before removing the dead bilinear lerp8 clamps (Codex: 161-node comb loop -> -0.042ns setup fail). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.1 KiB
Ch397 — 224139 LPDDR alpha integration plan
Established state
zsrt139a2is the current live 640x480 opaque board composite from the matching 224139 capture: six epochs, 98,163 written pixels, zero board drops.- The first in-bounds blended draw family is PSMT8 perspective triangles
(
idx31075+,TBP=13824,CBP=14282, 128x128,PRIM.ABE=1). - Its recorded
ALPHA_1state is0x0000008000000062; this is a generic selector/FIX blend, not the existing source-over-only shortcut.
Implemented cut (local RTL)
The current SH3 renderer sends final opaque pixels directly to
gs_lpddr_zc_emit. Its full-frame colour lives only in LPDDR, while the
existing combined triangle alpha path reads destination colour from BRAM.
Therefore a full-frame alpha implementation must extend the LPDDR emitter:
- carry
ABEplus ALPHA selector/FIX metadata with each fragment packet; - issue a single-pixel colour read after a Z-pass (or immediately for a non-Z alpha fragment);
- apply the existing
gs_alpha_blendgeneric selector equation in the emitter's AXI domain; - enqueue that blended colour through the existing colour writer, preserving the ordered scene-marker drain contract.
Implemented in the local RTL:
gs_lpddr_zc_emitnow carries{ABE,A,B,C,D,FIX}through its request and colour-align FIFOs;gs_lpddr_color_blendbypasses opaque pixels and, for ABE pixels, reads the containing 256-bit LPDDR beat, selects the addressed 32-bit lane, invokes genericgs_alpha_blend, and emits the ordered write payload;gs_lpddr_rd_arbhas a fifth client for that destination read, below scanout and Z/reload traffic;- the real raster wrapper forwards the per-primitive alpha snapshot.
make -C sim tb_gs_lpddr_color_blend passes the actual 0x62 selector
pattern: source 0x80402010 over destination 0x40203020 produces
0x80001010 (Cd - Cs, clamped), proving the LPDDR read lane and blend
datapath together.
The complete three-epoch integration replay is exact: 102,493 fragments,
102,479 Z passes, three ordered drains, zero drops/overflows/BRESP errors,
Z 0/307200 mismatches, and color 0/51525 mismatches. The final
640x480 simulation framebuffer is
sh3_zsrt139a3_sim_fb.mem (SHA-256
a331fd9d68aaccb35310051ad9091f6c02acafcc4b86391cecedaea8f08ce28f).
Timing closure and deploy artifact (2026-07-15)
The first alpha fit exposed the destination-read/blend arithmetic as an unregistered 310 MHz cone. The final implementation uses a registered shift/add blend engine (no new DSPs), a registered generic async-FIFO write port, and registered request-head control. The last 42 ps setup family was the staged scene-marker bit feeding the color-align RAM write enable; fragment valid and marker valid are now separate registered controls, removing that decode from color writes.
The final Quartus 25.3.1 build is signoff-clean: setup +0.012 ns, hold
0.000 ns, zero setup TNS, and synthesis/fit/STA/assembler all completed with
zero errors. The split loader artifact is
output_files/retroDE_ps2.core.rbf, SHA-256
f921f0bcde7fc9a8fe01bf39166327101e6dcbf0faf68398082207c8498ca349.
build_quartus.sh now explicitly runs post_flow.tcl after its direct
quartus_asm step, because direct assembler invocation does not dispatch the
QSF post-flow hook by itself.
The RBF and all zsrt139a3 host assets were copied to the DE25 and SHA-checked.
The first runtime load hit an SDM service-buffer timeout and left configfs in
an uninterruptible write. A controlled remote reboot restored network ping,
but SSH had not restarted at closeout time; a physical power cycle is the
remaining prerequisite before retrying the load and taking the board FB dump.
First silicon run and request-FIFO correction (2026-07-16)
The authoritative Quartus 26.1 GUI RBF loaded successfully and all three
texture uploads, readback CRCs, runtime palettes, staging counts, Z preclear,
and zero-drop checks passed. The render itself did not drain: epoch 0 reached
36,958 framebuffer beats after 120 seconds and frame_drained remained low.
The resulting interrupted composite was 298,888/307,200 pixels equal to the
simulation framebuffer, with 8,312 mismatches caused by rebinding texture and
palette state while old fragments were still executing.
Root cause was the Ch397 timing cut in generic gs_async_fifo: the allocation
pointer was published to the read clock when the write was accepted, one
writer cycle before the staged RAM write committed. The production request
FIFO crosses a 40 MHz writer into a roughly 310 MHz reader, so the consumer
could observe the pointer about 19 ns before the data existed and replay stale
RAM entries. The earlier FIFO regression used a faster writer than reader and
therefore did not exercise this CDC ordering failure.
The FIFO now maintains separate allocation and commit pointers. Full
accounting uses allocation, while only the pointer advanced by an actual
staged RAM commit crosses into the reader. A production-ratio stress test
(40 MHz write, 312.5 MHz read) transfers 3,284/3,284 ordered words with zero
duplicates/drops. The full zsrt139a3 integration replay remains exact:
102,493 fragments, all three drains, zero drops/errors, and Z/color scoreboards
at zero mismatches. This correction is ready for the owner-controlled
Quartus 26.1 GUI compile and subsequent silicon rerun.
Corrected FIFO silicon acceptance (2026-07-16)
The owner-controlled Quartus 26.1 GUI compile completed successfully. The
final fit used 37,153/46,800 ALMs and 284/358 RAM blocks; setup slack was
+0.064 ns, hold slack was 0.000 ns, and every reported TNS was zero. The
fresh split RBF is output_files/retroDE_ps2.core.rbf, 3,915,776 bytes,
SHA-256 a1adf0a706465709e3078dedc053433e34807a40a57733c1cd2e14c9d64461fa.
Its hash was checked again on the DE25 before core_loader.sh loaded it, and
fpga0 reported operating.
The corrected request FIFO is now proven on silicon. Epochs 0, 1, and 2
completed with 54, 68, and 10 records respectively; all texture CRC, runtime
palette, staging, Z-preclear, and zero-drop gates passed. The 10-record alpha
epoch is short enough for frame_drained to clear and rise entirely between
HPS bridge polls, so the host now accepts that missed-low case only if both the
new per-GO record count and framebuffer-beat count prove forward progress. A
one-second low-observation window replaces the old 120-second delay; the
full board run now completes in about 23 seconds with rc=0.
The final board framebuffer and simulation framebuffer are byte-for-byte
identical across all 307,200 PSMCT32 words (zero mismatches), both with
SHA-256 a331fd9d68aaccb35310051ad9091f6c02acafcc4b86391cecedaea8f08ce28f.
First acceptance scope
The generated zsrt139a3 fixture carries two opaque setup epochs followed by
authentic blended draw 32391 (10 triangles, safely within the 2048-word
staging limit). The six-draw family remains the next scale-up target; its
first four draws exceed that current staging window and must be split across
ordered epochs before board replay. Do not claim frame fidelity until the
subset is locally replayed and board-read back against its own trace.