Snapshot: fog implementation + fidelity tooling baseline (pre bilinear-clamp fix)
Per-vertex GS fog end-to-end (gs_stub emit incl. persp_emit5, gs_prim_list_feeder XYZ2->XYZF2 on PRIM.FGE, gs_make_sh3_scheduler_fixture.py F/FGE packing), new fog TBs, fidelity attribution tooling. Functional baseline before removing the dead bilinear lerp8 clamps (Codex: 161-node comb loop -> -0.042ns setup fail). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,151 @@
|
||||
# Ch437 — zero-line-RAM horizontal-linear presentation prefit
|
||||
|
||||
## Purpose
|
||||
|
||||
Ch436 made the captured 15:14 vertical presentation linear while retaining the
|
||||
Ch418 nearest-neighbour 5:4 horizontal map. The accepted framebuffer and board
|
||||
replay are exact, so this chapter changes presentation only: it reconstructs
|
||||
the fractional horizontal samples that nearest presentation repeats.
|
||||
|
||||
For output column `x`, the exact source coordinate is:
|
||||
|
||||
- `x0 = floor(x*4/5)`;
|
||||
- `xf = (x*4) mod 5`;
|
||||
- `x1 = min(x0+1, 511)`.
|
||||
|
||||
The existing five-state mapper already produces both `x0` and `xf`; no divider
|
||||
or new wide coordinate cone is introduced.
|
||||
|
||||
## Architecture
|
||||
|
||||
The rejected Ch436 horizontal prototype mirrored both complete line buffers to
|
||||
obtain an `x+1` read port. Quartus implemented the mirrors as 30,720 registers
|
||||
and rejected the design at 6,270 required LABs. Ch437 does not restore those
|
||||
arrays.
|
||||
|
||||
Instead, each physical line buffer retains one video read address. Two
|
||||
256-bit register slots hold the current and next beat for each of the two
|
||||
resident source rows. Horizontal blanking alternately primes beats 0 and 1.
|
||||
When active scanout enters beat K, the retired parity slot is refilled with
|
||||
beat K+1. The 5:4 cadence provides ten output clocks for each eight-source-pixel
|
||||
beat, so the next beat is resident well before the single lane-7 boundary
|
||||
sample needs it.
|
||||
|
||||
The interpolation order is separable and matches the software oracle:
|
||||
|
||||
1. blend row L and L+1 at x0 with denominator 15;
|
||||
2. blend row L and L+1 at x1 with denominator 15;
|
||||
3. blend those two results with denominator 5.
|
||||
|
||||
The source-right edge clamps at pixel 511. Other profiles retain their existing
|
||||
nearest/1:1 path through default-off parameters.
|
||||
|
||||
## Evidence before owner fit
|
||||
|
||||
Directed and legacy regressions:
|
||||
|
||||
- `tb_gs_lpddr_scanout_lb_hstretch`: 2,400 checked, zero errors, zero underflow,
|
||||
across eight 256-bit beats and the clamped right edge;
|
||||
- `tb_gs_lpddr_scanout_lb_psm32`: 49,152 checked, zero errors/underflow;
|
||||
- `tb_gs_lpddr_scanout_lb_psm32_256`: 196,608 checked, zero errors/underflow;
|
||||
- `tb_gs_lpddr_scanout_fb`: 105,968 checked, zero errors/underflow;
|
||||
- Verilator lint: no errors;
|
||||
- relevant `git diff --check`: pass.
|
||||
|
||||
The exact software presentation model improves the accepted f52 framebuffer's
|
||||
PCSX2 comparison again:
|
||||
|
||||
- Ch436 vertical-linear: MAE/RMSE `14.2678 / 21.3085`;
|
||||
- Ch437 separable-linear: MAE/RMSE `13.9525 / 20.6568`.
|
||||
|
||||
This is a visible smoothing of the horizontal staircase/noise without changing
|
||||
framebuffer bytes, render ordering, LPDDR traffic, or draw count.
|
||||
|
||||
## Owner GUI compile acceptance
|
||||
|
||||
### Rejected first fit
|
||||
|
||||
The first owner fit rejected the initial cache implementation at placement:
|
||||
6,337 LABs were required on the 4,680-LAB device. The synthesis/netlist report,
|
||||
not the summary alone, showed the exact cause:
|
||||
|
||||
- `u_lpddr_scan_lb` contained 32,312 registers;
|
||||
- the SDC line-buffer reset collection found 30,720 destinations;
|
||||
- `lb0[0..79]` and `lb1[0..79]` appeared as individual flip-flop nodes;
|
||||
- no line-buffer M20K payload remained in the scanout hierarchy.
|
||||
|
||||
The conditional reads used to prime/refill the cache prevented Quartus from
|
||||
recognizing the dual-clock RAM template. This was the same physical failure
|
||||
mode as the rejected Ch436 mirrors, despite the source-level arrays not being
|
||||
duplicated.
|
||||
|
||||
### Corrected inference-safe candidate
|
||||
|
||||
Each physical array now has exactly one unconditional video-clock read:
|
||||
|
||||
```systemverilog
|
||||
lb0_video_rd_q <= lb0[video_rd_addr_q];
|
||||
lb1_video_rd_q <= lb1[video_rd_addr_q];
|
||||
```
|
||||
|
||||
Address, valid, and beat tag are registered beside those outputs. A separate
|
||||
control process fills the two-beat register cache only from the registered
|
||||
RAM outputs. It never indexes `lb0` or `lb1` directly. This is the canonical
|
||||
simple-dual-port, dual-clock inference form already proven by the accepted
|
||||
line-buffer implementation.
|
||||
|
||||
The corrected candidate retains identical horizontal results: the expanded
|
||||
2,400-pixel test, all PSMCT32 legacy regressions, Verilator lint, and relevant
|
||||
`git diff --check` pass again.
|
||||
|
||||
The corrected RTL was accepted by an owner-controlled Quartus 26.1 GUI compile.
|
||||
The generated core is:
|
||||
|
||||
- `retroDE_ps2.core.rbf`, 4,132,864 bytes;
|
||||
- SHA-256
|
||||
`3dd624cd34e56cf553fc92c47c3a0d47dc6a9c0513f37cbf72f2c4f6f9c71559`.
|
||||
|
||||
All report gates pass:
|
||||
|
||||
- fitter successful at 39,077 / 46,800 ALMs (83%), 56,526 registers,
|
||||
5,327,792 block-memory bits, 344 / 358 RAM blocks, and 163 / 376 DSPs;
|
||||
- both line buffers are inferred as simple-dual-port, dual-clock
|
||||
`altera_syncram` instances, five M20Ks and 15,360 bits apiece;
|
||||
- the scanout hierarchy contains 1,602 registers rather than the rejected
|
||||
implementation's 32,312;
|
||||
- EMIF setup slack is +0.069 ns with zero TNS;
|
||||
- the 25 MHz design-clock setup slack is +5.473 ns with zero TNS;
|
||||
- all hold classes have zero TNS and minimum slack 0.000 ns;
|
||||
- minimum recovery, removal, and pulse-width slack are respectively
|
||||
+1.056 ns, +0.164 ns, and +0.200 ns.
|
||||
|
||||
## Hardware acceptance
|
||||
|
||||
The accepted core was installed on the DE25-Nano and loaded successfully.
|
||||
The live fabric reported the expected core ID/ABI, healthy status, active HDMI
|
||||
scanout, an advancing frame counter, and zero raster overflow.
|
||||
|
||||
The complete f52 scheduler replay then passed on hardware:
|
||||
|
||||
- all 311 epochs completed;
|
||||
- all fresh texture CRCs matched and resident reuse checks passed;
|
||||
- every epoch reported zero fragment drops;
|
||||
- `DONE rc=0 (all gates passed)`;
|
||||
- the dump contained 307,200 words, `sum32=0xaad0b94d`,
|
||||
`xor32=0x33138181`, 245,155 nonzero pixels, and bounds `(0,0)..(511,479)`.
|
||||
|
||||
The retrieved Ch437 board framebuffer has SHA-256
|
||||
`d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b`
|
||||
and is byte-identical to the accepted Ch436 framebuffer. This is the intended
|
||||
result: Ch437 changes only the HDMI presentation filter, after framebuffer
|
||||
storage, while preserving the proven renderer and replay output.
|
||||
|
||||
Ch437 is accepted. Its closure criteria were:
|
||||
|
||||
- fitter completes within the device's 4,680-LAB capacity;
|
||||
- line buffers remain inferred M20Ks and total RAM stays at 344 / 358 blocks;
|
||||
- all setup, hold, recovery, removal, and pulse-width classes are nonnegative;
|
||||
- the narrow EMIF request-address family remains clean;
|
||||
- the 25 MHz presentation domain remains clean.
|
||||
|
||||
All five criteria passed before deployment.
|
||||
Reference in New Issue
Block a user