Files
retroDE_ps2/docs/ch439_timing_convergence_prefit.md
T
thejayman77 ba74bbd5aa Snapshot: fog implementation + fidelity tooling baseline (pre bilinear-clamp fix)
Per-vertex GS fog end-to-end (gs_stub emit incl. persp_emit5, gs_prim_list_feeder
XYZ2->XYZF2 on PRIM.FGE, gs_make_sh3_scheduler_fixture.py F/FGE packing), new fog
TBs, fidelity attribution tooling. Functional baseline before removing the dead
bilinear lerp8 clamps (Codex: 161-node comb loop -> -0.042ns setup fail).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 19:56:46 -04:00

280 lines
15 KiB
Markdown

# Ch439 — measured FIFO right-sizing and scanout write-port timing convergence
## Why this chapter exists
The second Ch438 GUI fit completed successfully but was not timing-clean and
was not deployed. Its RBF SHA-256 was
`b5865b599101d429b86c67c25ae9688c0f6dda334aa7ebbcbd664f31226aa602`.
The fit used 39,442 ALMs, 56,446 registers, 349/358 M20Ks, and 160 DSPs.
The actual post-fit STA report showed three concrete families:
- setup WNS -0.924 ns / TNS -12.931 ns in the 310 MHz EMIF domain;
- six of the top ten setup paths were
`u_zc_emit|u_req|raddr_lo_q[1]` to physical `mem_lo` RAM address ports;
- one setup path at -0.403 ns was `u_lpddr_scan_lb|r_data_q[240]` to the
`lb1` RAM write port;
- the only negative design-clock hold path was -4.221 ns from the bridge's
quasi-static `lpddr_video_src_q` to `vsrc_cap_d` across unrelated clocks.
The preceding two-way width bank did not reduce the request RAM count and made
placement worse than the original packed 16K FIFO. This chapter therefore
removes that banked form from the production instance instead of chasing it
with another seed.
## Structural correction
### Request FIFO: 16K to 8K from measured occupancy
The complete 311-epoch f52 production replay was rerun at depth 16,384 while
recording the per-epoch request occupancy. The maximum was 6,115 entries at
epoch 148. The next power-of-two depth, 8,192, therefore leaves 2,077 entries
or 34 percent measured headroom. It also halves the packed 93-bit request
FIFO's depth, expected M20K footprint, and physical read-address tree.
The production `u_zc_emit` instance now uses `REQ_DEPTH=8192` and the ordinary
monolithic registered-read FIFO. `zc_g_drops` remains the fail-closed runtime
guard if a future fixture exceeds the measured envelope.
### Scanout line buffers: RAM-local physical write stage
The response stage no longer drives all three line-buffer RAM write ports.
`L_R` copies each returned beat directly into data/address/enable registers
dedicated to the selected physical RAM. Those registers commit independently
while the FSM issues or waits for the next single-beat read; `L_C` is used only
once to flush the final beat of a row. The fitter can place each 256-bit launch
bank beside its RAM without adding a state to every beat.
### LPDDR video-source bundled CDC
`lpddr_video_src_q` is written before `lpddr_ctrl_commit_w` toggles.
`vsrc_cap_d` samples the stable data only after that toggle traverses the
three-flop `commit_d_sync`. The raw zero-cycle hold analysis between unrelated
clocks is therefore non-functional. The SDC now cuts only HOLD for this exact
source/destination pair and retains a 2 ns maximum net-delay bound. Setup is
not false-pathed.
## Verification before the next GUI compile
- binomial scanout: 2,400/2,400 exact pixels, zero underflow/read errors, PASS;
- horizontal-stretch scanout: 2,400/2,400 exact pixels, zero underflow/read
errors, PASS;
- registered asynchronous FIFO stress: 3,285 writes/reads, zero errors, PASS;
- directed Z/color integration: 1,500 fragments, zero color overflow/BRESP
errors, PASS;
- complete f52 replay at the shipping 8K request depth:
- 311 epochs;
- 2,339,914 fragments fed and 2,286,292 passed;
- zero drops, Z mismatches, color mismatches, color overflow, or BRESP errors;
- PASS;
- Ch439 framebuffer SHA-256:
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`;
- byte-identical to the accepted Ch437 framebuffer and the prior Ch438
production simulation.
No Ch438 timing-red RBF has been deployed. A fresh Quartus 26.1 GUI compile is
required for this consolidated Ch439 correction.
## First Ch439 fit and hardware feedback
The first owner GUI fit of the resource-relieved design was fully signoff-clean:
EMIF setup +0.098 ns, all hold classes nonnegative, 39,514 ALMs, 57,331
registers, 309/358 M20Ks, and 160 DSPs. It also proved that the 8K cut removed
exactly 40 RAM blocks as predicted. The generated RBF SHA-256 was
`7167deb651c08057828a360b6083ec27eb140d59eb08b7b06315cbc1d235ec13`.
It loaded successfully and the complete 311-epoch hardware replay passed every
renderer gate; its framebuffer was byte-identical to simulation at SHA-256
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`.
However, the HDMI line-buffer status was not acceptable. `LPDDR_STATUS` held
`0xF1`, with bit 5 set. Disabling LPDDR scanout cleared it to `0xD1`, and
re-enabling scanout restored `0xF1`, proving sustained line-buffer underflow
rather than a sticky AXI read error or startup-only event.
Root cause was the first Ch439 `L_R -> L_C -> L_W` write pipeline, which added
one EMIF cycle to every 32-byte beat and crossed the real just-in-time refill
margin. Ch439b removed `L_W`, but board testing proved that retaining `L_C`
after every beat was still too slow. Its clean fit used 39,616 ALMs, 57,109
registers, 309/358 M20Ks, and 160 DSPs; EMIF setup was +0.050 ns and every
other timing class was also clean. RBF SHA-256 was
`56e45346338907eb6c307320757774ac65efa56dabb644906a5600f5ea73a982`.
The full 311-epoch board replay again passed every renderer gate and produced
the exact expected framebuffer, but four live samples across advancing frames
all read `LPDDR_STATUS=0xF1`. Source-off cleared status to `0xD1`; source-on
restored `0xF1` immediately. Ch439b is therefore rejected as well.
Its board evidence is retained as
`sim/data/top_psmct32_raster_demo/sh3_zsrt139f52_ch439b_board_run.log` and
`sh3_zsrt139f52_ch439b_board_fb.mem`; the framebuffer is byte-identical to the
8K production simulation at SHA-256
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`.
Ch439c turns the RAM-local write stage into an actual pipeline. A captured
beat commits from its local register while the AXI FSM advances to the next
read. Only the final beat takes `L_C`, as a one-cycle row-end flush before
`next_fetch` is published. This restores the accepted pre-Ch439 per-beat
cadence while retaining the physical timing cut. Directed binomial and
horizontal-stretch tests are exact at 2,400/2,400 pixels with zero underflow or
read errors; the 640x192 and 256x768 PSMCT32 tests are exact at 49,152/49,152
and 196,608/196,608 pixels with zero underflow or read errors. A fresh owner
GUI compile is required for Ch439c.
## Ch439c fit feedback and Ch439d address-tree cut
The owner Ch439c GUI compile completed successfully but was not deployed. It
used 39,472 ALMs, 56,640 registers, 309/358 M20Ks, and 160 DSPs. Hold,
recovery, removal, and minimum-pulse-width checks were clean, but EMIF setup
failed at -1.009 ns / -60.117 ns TNS across 166 endpoints. All ten reported
worst paths were the same zero-logic-level route:
- launch: `u_zc_emit|u_req|raddr_q[7]`;
- capture: the monolithic request FIFO's physical `mem_rtl_0` port-B address
registers;
- data delay: 4.077 ns, of which 3.759 ns was interconnect.
This is not a scanout-pipeline path. The Ch439c placement exposed the remaining
8K request-RAM address tree. Ch439d enables the already-regressed two-way
payload bank only on `u_zc_emit.u_req`. At the present 8K depth, each preserved
read-address copy drives half the width at half the original Ch438 depth: about
one quarter of the physical RAM load that defeated the earlier 16K banked fit.
Depth, ordering, registered-read latency, and the measured 2,077-entry capacity
margin are unchanged.
Post-cut verification is complete:
- banked asynchronous FIFO stress: 3,285 writes and reads, zero errors, PASS;
- directed Z/color integration: 1,500 fragments, zero overflow/BRESP errors,
PASS;
- complete 311-epoch f52 replay at 8K: 2,339,914 fragments fed, 2,286,292
passed, zero drops/mismatches/errors, PASS;
- framebuffer remains byte-identical at SHA-256
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`.
The Ch439c timing-red RBF SHA-256 is
`d006cc1ede2db1eeb5b8ae961eafc047eff3e6ef3ce58a173ae4f8e3bf7c3cac` and
was not deployed. A fresh owner GUI compile is required for Ch439d.
Board evidence is retained as
`sim/data/top_psmct32_raster_demo/sh3_zsrt139f52_ch439_board_run.log` and
`sh3_zsrt139f52_ch439_board_fb.mem`. The HDMI-equivalent filtered preview is
`captures/gs/silenthill3/extracted/recon/sh3_zsrt139f52_ch439_hdmi.png` and
remains MAE 13.1111 / RMSE 19.1051 against the PCSX2 reference.
## Ch439d fit feedback and Ch439e quadrant cut
The owner Ch439d GUI compile completed successfully and the two width banks
were inferred as intended, but the RBF was not deployed because EMIF setup
still failed. Its RBF SHA-256 is
`fe6eb909b8d875b617587f8afafdc510f77e4b5369c66c225ba9cea404ef6ba9`.
The actual post-fit results were:
- 39,410 required ALMs and 42,843 final-placement ALMs (91 percent);
- 56,822 dedicated logic registers, 309/358 M20Ks, and 160 DSPs;
- setup WNS -0.515 ns / TNS -36.152 ns, with 120 failing endpoints;
- hold, recovery, removal, and minimum-pulse-width checks all clean.
Every reported worst setup path remained one zero-logic-level physical-route
family from `u_req|raddr_{lo,hi}_q[*]` to the corresponding `mem_{lo,hi}` M20K
address input register. The worst path was `raddr_lo_q[10]` to a `mem_lo`
address register, with 3.615 ns data delay into the 3.225 ns EMIF period.
Width banking improved WNS by about 0.5 ns versus Ch439c but left each address
copy spanning the full 8K depth, so Ch439d is rejected rather than reseeded.
Reducing the FIFO below 8K is not valid for the current producer. The existing
instrumented complete f52 replay measured a peak request occupancy of 6,115 at
epoch 148, and the producer drops fragments when `zc_g_ready` deasserts. The
8,192-entry depth and its measured 2,077-entry margin therefore remain fixed.
Similarly, directly feeding the next binary pointer into the RAM address port
was already exercised before Ch420 and exposed pointer logic plus the same
high-fanout physical address family.
Ch439e instead crosses two depth banks with the existing two payload-width
banks. The request FIFO is physically four 4K-deep RAM quadrants:
`mem_lo0`, `mem_lo1`, `mem_hi0`, and `mem_hi1`. Each has its own preserved
11-bit RAM-facing read-address register, reducing one launch tree to roughly
one quarter of the original physical RAM load. A one-bit selector delayed with
the synchronous RAM data performs only a 2:1 lower/upper-depth selection at
the registered output. Capacity, ordering, CDC publication, accepted-read
latency, and atomic tuple writes are unchanged.
Ch439e prefit verification is complete:
- quadrant asynchronous FIFO stress: 3,285 writes and reads, zero errors,
PASS;
- directed Z/color integration: 1,500 fragments, zero overflow/BRESP errors,
PASS;
- complete 311-epoch f52 production replay at 8K: 2,339,914 fragments fed,
2,286,292 passed, zero drops/mismatches/errors, PASS;
- binomial and horizontal-stretch scanout: 2,400/2,400 exact pixels apiece,
zero underflow/read errors, PASS;
- framebuffer SHA-256 remains byte-identical at
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`.
The next owner GUI compile must confirm that all four quadrant memories and
four independent read-address trees survive synthesis, that setup closes, and
that the new RAM-output depth selector does not become a replacement setup
family. Only a fully signoff-clean Ch439e artifact is eligible for deployment.
## Ch439e fit feedback and Ch439f arithmetic-boundary correction
The owner Ch439e GUI compile completed successfully and confirmed the intended
request-FIFO result. All four quadrant memories were inferred, request-memory
usage stayed at 309/358 M20Ks, and the 310 MHz EMIF setup domain closed at
+0.019 ns. Hold, recovery, removal, and minimum-pulse-width checks were also
clean. The design used 38,849 required ALMs, 42,880 final-placement ALMs
(92 percent), 57,233 dedicated registers, and 160/376 DSP blocks.
The artifact is nevertheless rejected because the 25 MHz design clock failed
at WNS -1.294 ns / TNS -89.732 ns across 208 endpoints. Its rejected RBF
SHA-256 is
`4deb590b83036e250e042158896a3a69c731bbccc09eef04c06f523cc382c6e5`.
Every reported worst path was the same 27-logic-level, 41.141 ns arithmetic
family:
- launch: `u_demo|u_gs|u_tex|g_bilinear.tap[0][11]`;
- capture: `u_texcache|tex_mem...portbaddr[10]`;
- middle: both bilinear lerps, GS modulation/blend arithmetic, and the next
texture-coordinate/address arithmetic.
The initial hypothesis was that the project-wide `AUTO_RESOURCE_SHARING ON`
setting caused this apparently cross-transaction path. The fit report records
GS arithmetic operators packed with `u_tex` bilinear operators, including
`u_demo|u_gs|mult_66` merged into `u_demo|u_gs|u_tex|mult_9`; Intel's Quartus
26.1 setting reference also warns that sharing mux/control logic can reduce
Fmax. Ch439f tests that hypothesis directly below rather than assuming it.
Ch439f disabled automatic resource sharing globally while retaining
`AGGRESSIVE AREA`. This was a controlled synthesis experiment, not a seed or
RTL-latency change.
The fresh owner Ch439f compile conclusively rejected that hypothesis.
Synthesis reported `Auto Resource Sharing: Off`, but the complete result was
bit-for-bit structurally unchanged at report precision: 38,849 required ALMs,
42,880 final-placement ALMs, 57,233 registers, 309 M20Ks, 160 DSPs, EMIF setup
+0.019 ns, and the identical design-clock -1.294 ns / -89.732 ns / 208-endpoint
tap-to-texture-cache family with the same 41.141 ns path. The project setting
is therefore restored to its area-oriented value.
## Ch439g functional exception for the impossible packed-DSP arc
The RTL cone proves that a captured bilinear tap cannot reach the texture read
address combinationally. `tap[*]` feeds the two lerp levels only in `BS_DONE`,
where the result is captured into `tex_color_hold`; the external filtered
result is exposed from that register only in the distinct `BS_OUT` cycle.
Texture-memory address capture occurs under `tex_rd_en` in `BS_ISSUE` or a
nearest-read issue cycle. Those state conditions are mutually exclusive.
The remaining path is therefore a conservative physical arc introduced by
packing mutually exclusive arithmetic into Agilex DSP structures, not a
single-cycle functional requirement. Ch439g adds a fail-closed SDC exception
from exactly the architectural 128 `g_bilinear.tap[0:3][31:0]` registers to
the existing `u_texcache.tex_mem` endpoint collection. The exception does not
cut any walker, affine-UV, perspective-UV, descriptor, or address register;
all real texture-address launches remain timed normally. The SDC halts the
compile if the tap population is not exactly 128 or the cache collection is
empty.
No functional RTL changed after the already-exact Ch439e simulations. The next
owner GUI compile must show the Ch439g collection-count message, retain the
four request quadrants and EMIF closure, and expose the true next design-clock
family. Deployment remains forbidden unless every timing class is clean.