# Ch439 — measured FIFO right-sizing and scanout write-port timing convergence ## Why this chapter exists The second Ch438 GUI fit completed successfully but was not timing-clean and was not deployed. Its RBF SHA-256 was `b5865b599101d429b86c67c25ae9688c0f6dda334aa7ebbcbd664f31226aa602`. The fit used 39,442 ALMs, 56,446 registers, 349/358 M20Ks, and 160 DSPs. The actual post-fit STA report showed three concrete families: - setup WNS -0.924 ns / TNS -12.931 ns in the 310 MHz EMIF domain; - six of the top ten setup paths were `u_zc_emit|u_req|raddr_lo_q[1]` to physical `mem_lo` RAM address ports; - one setup path at -0.403 ns was `u_lpddr_scan_lb|r_data_q[240]` to the `lb1` RAM write port; - the only negative design-clock hold path was -4.221 ns from the bridge's quasi-static `lpddr_video_src_q` to `vsrc_cap_d` across unrelated clocks. The preceding two-way width bank did not reduce the request RAM count and made placement worse than the original packed 16K FIFO. This chapter therefore removes that banked form from the production instance instead of chasing it with another seed. ## Structural correction ### Request FIFO: 16K to 8K from measured occupancy The complete 311-epoch f52 production replay was rerun at depth 16,384 while recording the per-epoch request occupancy. The maximum was 6,115 entries at epoch 148. The next power-of-two depth, 8,192, therefore leaves 2,077 entries or 34 percent measured headroom. It also halves the packed 93-bit request FIFO's depth, expected M20K footprint, and physical read-address tree. The production `u_zc_emit` instance now uses `REQ_DEPTH=8192` and the ordinary monolithic registered-read FIFO. `zc_g_drops` remains the fail-closed runtime guard if a future fixture exceeds the measured envelope. ### Scanout line buffers: RAM-local physical write stage The response stage no longer drives all three line-buffer RAM write ports. `L_R` copies each returned beat directly into data/address/enable registers dedicated to the selected physical RAM. Those registers commit independently while the FSM issues or waits for the next single-beat read; `L_C` is used only once to flush the final beat of a row. The fitter can place each 256-bit launch bank beside its RAM without adding a state to every beat. ### LPDDR video-source bundled CDC `lpddr_video_src_q` is written before `lpddr_ctrl_commit_w` toggles. `vsrc_cap_d` samples the stable data only after that toggle traverses the three-flop `commit_d_sync`. The raw zero-cycle hold analysis between unrelated clocks is therefore non-functional. The SDC now cuts only HOLD for this exact source/destination pair and retains a 2 ns maximum net-delay bound. Setup is not false-pathed. ## Verification before the next GUI compile - binomial scanout: 2,400/2,400 exact pixels, zero underflow/read errors, PASS; - horizontal-stretch scanout: 2,400/2,400 exact pixels, zero underflow/read errors, PASS; - registered asynchronous FIFO stress: 3,285 writes/reads, zero errors, PASS; - directed Z/color integration: 1,500 fragments, zero color overflow/BRESP errors, PASS; - complete f52 replay at the shipping 8K request depth: - 311 epochs; - 2,339,914 fragments fed and 2,286,292 passed; - zero drops, Z mismatches, color mismatches, color overflow, or BRESP errors; - PASS; - Ch439 framebuffer SHA-256: `d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`; - byte-identical to the accepted Ch437 framebuffer and the prior Ch438 production simulation. No Ch438 timing-red RBF has been deployed. A fresh Quartus 26.1 GUI compile is required for this consolidated Ch439 correction. ## First Ch439 fit and hardware feedback The first owner GUI fit of the resource-relieved design was fully signoff-clean: EMIF setup +0.098 ns, all hold classes nonnegative, 39,514 ALMs, 57,331 registers, 309/358 M20Ks, and 160 DSPs. It also proved that the 8K cut removed exactly 40 RAM blocks as predicted. The generated RBF SHA-256 was `7167deb651c08057828a360b6083ec27eb140d59eb08b7b06315cbc1d235ec13`. It loaded successfully and the complete 311-epoch hardware replay passed every renderer gate; its framebuffer was byte-identical to simulation at SHA-256 `d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`. However, the HDMI line-buffer status was not acceptable. `LPDDR_STATUS` held `0xF1`, with bit 5 set. Disabling LPDDR scanout cleared it to `0xD1`, and re-enabling scanout restored `0xF1`, proving sustained line-buffer underflow rather than a sticky AXI read error or startup-only event. Root cause was the first Ch439 `L_R -> L_C -> L_W` write pipeline, which added one EMIF cycle to every 32-byte beat and crossed the real just-in-time refill margin. Ch439b removed `L_W`, but board testing proved that retaining `L_C` after every beat was still too slow. Its clean fit used 39,616 ALMs, 57,109 registers, 309/358 M20Ks, and 160 DSPs; EMIF setup was +0.050 ns and every other timing class was also clean. RBF SHA-256 was `56e45346338907eb6c307320757774ac65efa56dabb644906a5600f5ea73a982`. The full 311-epoch board replay again passed every renderer gate and produced the exact expected framebuffer, but four live samples across advancing frames all read `LPDDR_STATUS=0xF1`. Source-off cleared status to `0xD1`; source-on restored `0xF1` immediately. Ch439b is therefore rejected as well. Its board evidence is retained as `sim/data/top_psmct32_raster_demo/sh3_zsrt139f52_ch439b_board_run.log` and `sh3_zsrt139f52_ch439b_board_fb.mem`; the framebuffer is byte-identical to the 8K production simulation at SHA-256 `d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`. Ch439c turns the RAM-local write stage into an actual pipeline. A captured beat commits from its local register while the AXI FSM advances to the next read. Only the final beat takes `L_C`, as a one-cycle row-end flush before `next_fetch` is published. This restores the accepted pre-Ch439 per-beat cadence while retaining the physical timing cut. Directed binomial and horizontal-stretch tests are exact at 2,400/2,400 pixels with zero underflow or read errors; the 640x192 and 256x768 PSMCT32 tests are exact at 49,152/49,152 and 196,608/196,608 pixels with zero underflow or read errors. A fresh owner GUI compile is required for Ch439c. ## Ch439c fit feedback and Ch439d address-tree cut The owner Ch439c GUI compile completed successfully but was not deployed. It used 39,472 ALMs, 56,640 registers, 309/358 M20Ks, and 160 DSPs. Hold, recovery, removal, and minimum-pulse-width checks were clean, but EMIF setup failed at -1.009 ns / -60.117 ns TNS across 166 endpoints. All ten reported worst paths were the same zero-logic-level route: - launch: `u_zc_emit|u_req|raddr_q[7]`; - capture: the monolithic request FIFO's physical `mem_rtl_0` port-B address registers; - data delay: 4.077 ns, of which 3.759 ns was interconnect. This is not a scanout-pipeline path. The Ch439c placement exposed the remaining 8K request-RAM address tree. Ch439d enables the already-regressed two-way payload bank only on `u_zc_emit.u_req`. At the present 8K depth, each preserved read-address copy drives half the width at half the original Ch438 depth: about one quarter of the physical RAM load that defeated the earlier 16K banked fit. Depth, ordering, registered-read latency, and the measured 2,077-entry capacity margin are unchanged. Post-cut verification is complete: - banked asynchronous FIFO stress: 3,285 writes and reads, zero errors, PASS; - directed Z/color integration: 1,500 fragments, zero overflow/BRESP errors, PASS; - complete 311-epoch f52 replay at 8K: 2,339,914 fragments fed, 2,286,292 passed, zero drops/mismatches/errors, PASS; - framebuffer remains byte-identical at SHA-256 `d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`. The Ch439c timing-red RBF SHA-256 is `d006cc1ede2db1eeb5b8ae961eafc047eff3e6ef3ce58a173ae4f8e3bf7c3cac` and was not deployed. A fresh owner GUI compile is required for Ch439d. Board evidence is retained as `sim/data/top_psmct32_raster_demo/sh3_zsrt139f52_ch439_board_run.log` and `sh3_zsrt139f52_ch439_board_fb.mem`. The HDMI-equivalent filtered preview is `captures/gs/silenthill3/extracted/recon/sh3_zsrt139f52_ch439_hdmi.png` and remains MAE 13.1111 / RMSE 19.1051 against the PCSX2 reference. ## Ch439d fit feedback and Ch439e quadrant cut The owner Ch439d GUI compile completed successfully and the two width banks were inferred as intended, but the RBF was not deployed because EMIF setup still failed. Its RBF SHA-256 is `fe6eb909b8d875b617587f8afafdc510f77e4b5369c66c225ba9cea404ef6ba9`. The actual post-fit results were: - 39,410 required ALMs and 42,843 final-placement ALMs (91 percent); - 56,822 dedicated logic registers, 309/358 M20Ks, and 160 DSPs; - setup WNS -0.515 ns / TNS -36.152 ns, with 120 failing endpoints; - hold, recovery, removal, and minimum-pulse-width checks all clean. Every reported worst setup path remained one zero-logic-level physical-route family from `u_req|raddr_{lo,hi}_q[*]` to the corresponding `mem_{lo,hi}` M20K address input register. The worst path was `raddr_lo_q[10]` to a `mem_lo` address register, with 3.615 ns data delay into the 3.225 ns EMIF period. Width banking improved WNS by about 0.5 ns versus Ch439c but left each address copy spanning the full 8K depth, so Ch439d is rejected rather than reseeded. Reducing the FIFO below 8K is not valid for the current producer. The existing instrumented complete f52 replay measured a peak request occupancy of 6,115 at epoch 148, and the producer drops fragments when `zc_g_ready` deasserts. The 8,192-entry depth and its measured 2,077-entry margin therefore remain fixed. Similarly, directly feeding the next binary pointer into the RAM address port was already exercised before Ch420 and exposed pointer logic plus the same high-fanout physical address family. Ch439e instead crosses two depth banks with the existing two payload-width banks. The request FIFO is physically four 4K-deep RAM quadrants: `mem_lo0`, `mem_lo1`, `mem_hi0`, and `mem_hi1`. Each has its own preserved 11-bit RAM-facing read-address register, reducing one launch tree to roughly one quarter of the original physical RAM load. A one-bit selector delayed with the synchronous RAM data performs only a 2:1 lower/upper-depth selection at the registered output. Capacity, ordering, CDC publication, accepted-read latency, and atomic tuple writes are unchanged. Ch439e prefit verification is complete: - quadrant asynchronous FIFO stress: 3,285 writes and reads, zero errors, PASS; - directed Z/color integration: 1,500 fragments, zero overflow/BRESP errors, PASS; - complete 311-epoch f52 production replay at 8K: 2,339,914 fragments fed, 2,286,292 passed, zero drops/mismatches/errors, PASS; - binomial and horizontal-stretch scanout: 2,400/2,400 exact pixels apiece, zero underflow/read errors, PASS; - framebuffer SHA-256 remains byte-identical at `d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`. The next owner GUI compile must confirm that all four quadrant memories and four independent read-address trees survive synthesis, that setup closes, and that the new RAM-output depth selector does not become a replacement setup family. Only a fully signoff-clean Ch439e artifact is eligible for deployment. ## Ch439e fit feedback and Ch439f arithmetic-boundary correction The owner Ch439e GUI compile completed successfully and confirmed the intended request-FIFO result. All four quadrant memories were inferred, request-memory usage stayed at 309/358 M20Ks, and the 310 MHz EMIF setup domain closed at +0.019 ns. Hold, recovery, removal, and minimum-pulse-width checks were also clean. The design used 38,849 required ALMs, 42,880 final-placement ALMs (92 percent), 57,233 dedicated registers, and 160/376 DSP blocks. The artifact is nevertheless rejected because the 25 MHz design clock failed at WNS -1.294 ns / TNS -89.732 ns across 208 endpoints. Its rejected RBF SHA-256 is `4deb590b83036e250e042158896a3a69c731bbccc09eef04c06f523cc382c6e5`. Every reported worst path was the same 27-logic-level, 41.141 ns arithmetic family: - launch: `u_demo|u_gs|u_tex|g_bilinear.tap[0][11]`; - capture: `u_texcache|tex_mem...portbaddr[10]`; - middle: both bilinear lerps, GS modulation/blend arithmetic, and the next texture-coordinate/address arithmetic. The initial hypothesis was that the project-wide `AUTO_RESOURCE_SHARING ON` setting caused this apparently cross-transaction path. The fit report records GS arithmetic operators packed with `u_tex` bilinear operators, including `u_demo|u_gs|mult_66` merged into `u_demo|u_gs|u_tex|mult_9`; Intel's Quartus 26.1 setting reference also warns that sharing mux/control logic can reduce Fmax. Ch439f tests that hypothesis directly below rather than assuming it. Ch439f disabled automatic resource sharing globally while retaining `AGGRESSIVE AREA`. This was a controlled synthesis experiment, not a seed or RTL-latency change. The fresh owner Ch439f compile conclusively rejected that hypothesis. Synthesis reported `Auto Resource Sharing: Off`, but the complete result was bit-for-bit structurally unchanged at report precision: 38,849 required ALMs, 42,880 final-placement ALMs, 57,233 registers, 309 M20Ks, 160 DSPs, EMIF setup +0.019 ns, and the identical design-clock -1.294 ns / -89.732 ns / 208-endpoint tap-to-texture-cache family with the same 41.141 ns path. The project setting is therefore restored to its area-oriented value. ## Ch439g functional exception for the impossible packed-DSP arc The RTL cone proves that a captured bilinear tap cannot reach the texture read address combinationally. `tap[*]` feeds the two lerp levels only in `BS_DONE`, where the result is captured into `tex_color_hold`; the external filtered result is exposed from that register only in the distinct `BS_OUT` cycle. Texture-memory address capture occurs under `tex_rd_en` in `BS_ISSUE` or a nearest-read issue cycle. Those state conditions are mutually exclusive. The remaining path is therefore a conservative physical arc introduced by packing mutually exclusive arithmetic into Agilex DSP structures, not a single-cycle functional requirement. Ch439g adds a fail-closed SDC exception from exactly the architectural 128 `g_bilinear.tap[0:3][31:0]` registers to the existing `u_texcache.tex_mem` endpoint collection. The exception does not cut any walker, affine-UV, perspective-UV, descriptor, or address register; all real texture-address launches remain timed normally. The SDC halts the compile if the tap population is not exactly 128 or the cache collection is empty. No functional RTL changed after the already-exact Ch439e simulations. The next owner GUI compile must show the Ch439g collection-count message, retain the four request quadrants and EMIF closure, and expose the true next design-clock family. Deployment remains forbidden unless every timing class is clean.