Files
retroDE_ps2/docs/ch437_horizontal_linear_prefit.md
T
thejayman77 ba74bbd5aa Snapshot: fog implementation + fidelity tooling baseline (pre bilinear-clamp fix)
Per-vertex GS fog end-to-end (gs_stub emit incl. persp_emit5, gs_prim_list_feeder
XYZ2->XYZF2 on PRIM.FGE, gs_make_sh3_scheduler_fixture.py F/FGE packing), new fog
TBs, fidelity attribution tooling. Functional baseline before removing the dead
bilinear lerp8 clamps (Codex: 161-node comb loop -> -0.042ns setup fail).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 19:56:46 -04:00

6.2 KiB

Ch437 — zero-line-RAM horizontal-linear presentation prefit

Purpose

Ch436 made the captured 15:14 vertical presentation linear while retaining the Ch418 nearest-neighbour 5:4 horizontal map. The accepted framebuffer and board replay are exact, so this chapter changes presentation only: it reconstructs the fractional horizontal samples that nearest presentation repeats.

For output column x, the exact source coordinate is:

  • x0 = floor(x*4/5);
  • xf = (x*4) mod 5;
  • x1 = min(x0+1, 511).

The existing five-state mapper already produces both x0 and xf; no divider or new wide coordinate cone is introduced.

Architecture

The rejected Ch436 horizontal prototype mirrored both complete line buffers to obtain an x+1 read port. Quartus implemented the mirrors as 30,720 registers and rejected the design at 6,270 required LABs. Ch437 does not restore those arrays.

Instead, each physical line buffer retains one video read address. Two 256-bit register slots hold the current and next beat for each of the two resident source rows. Horizontal blanking alternately primes beats 0 and 1. When active scanout enters beat K, the retired parity slot is refilled with beat K+1. The 5:4 cadence provides ten output clocks for each eight-source-pixel beat, so the next beat is resident well before the single lane-7 boundary sample needs it.

The interpolation order is separable and matches the software oracle:

  1. blend row L and L+1 at x0 with denominator 15;
  2. blend row L and L+1 at x1 with denominator 15;
  3. blend those two results with denominator 5.

The source-right edge clamps at pixel 511. Other profiles retain their existing nearest/1:1 path through default-off parameters.

Evidence before owner fit

Directed and legacy regressions:

  • tb_gs_lpddr_scanout_lb_hstretch: 2,400 checked, zero errors, zero underflow, across eight 256-bit beats and the clamped right edge;
  • tb_gs_lpddr_scanout_lb_psm32: 49,152 checked, zero errors/underflow;
  • tb_gs_lpddr_scanout_lb_psm32_256: 196,608 checked, zero errors/underflow;
  • tb_gs_lpddr_scanout_fb: 105,968 checked, zero errors/underflow;
  • Verilator lint: no errors;
  • relevant git diff --check: pass.

The exact software presentation model improves the accepted f52 framebuffer's PCSX2 comparison again:

  • Ch436 vertical-linear: MAE/RMSE 14.2678 / 21.3085;
  • Ch437 separable-linear: MAE/RMSE 13.9525 / 20.6568.

This is a visible smoothing of the horizontal staircase/noise without changing framebuffer bytes, render ordering, LPDDR traffic, or draw count.

Owner GUI compile acceptance

Rejected first fit

The first owner fit rejected the initial cache implementation at placement: 6,337 LABs were required on the 4,680-LAB device. The synthesis/netlist report, not the summary alone, showed the exact cause:

  • u_lpddr_scan_lb contained 32,312 registers;
  • the SDC line-buffer reset collection found 30,720 destinations;
  • lb0[0..79] and lb1[0..79] appeared as individual flip-flop nodes;
  • no line-buffer M20K payload remained in the scanout hierarchy.

The conditional reads used to prime/refill the cache prevented Quartus from recognizing the dual-clock RAM template. This was the same physical failure mode as the rejected Ch436 mirrors, despite the source-level arrays not being duplicated.

Corrected inference-safe candidate

Each physical array now has exactly one unconditional video-clock read:

lb0_video_rd_q <= lb0[video_rd_addr_q];
lb1_video_rd_q <= lb1[video_rd_addr_q];

Address, valid, and beat tag are registered beside those outputs. A separate control process fills the two-beat register cache only from the registered RAM outputs. It never indexes lb0 or lb1 directly. This is the canonical simple-dual-port, dual-clock inference form already proven by the accepted line-buffer implementation.

The corrected candidate retains identical horizontal results: the expanded 2,400-pixel test, all PSMCT32 legacy regressions, Verilator lint, and relevant git diff --check pass again.

The corrected RTL was accepted by an owner-controlled Quartus 26.1 GUI compile. The generated core is:

  • retroDE_ps2.core.rbf, 4,132,864 bytes;
  • SHA-256 3dd624cd34e56cf553fc92c47c3a0d47dc6a9c0513f37cbf72f2c4f6f9c71559.

All report gates pass:

  • fitter successful at 39,077 / 46,800 ALMs (83%), 56,526 registers, 5,327,792 block-memory bits, 344 / 358 RAM blocks, and 163 / 376 DSPs;
  • both line buffers are inferred as simple-dual-port, dual-clock altera_syncram instances, five M20Ks and 15,360 bits apiece;
  • the scanout hierarchy contains 1,602 registers rather than the rejected implementation's 32,312;
  • EMIF setup slack is +0.069 ns with zero TNS;
  • the 25 MHz design-clock setup slack is +5.473 ns with zero TNS;
  • all hold classes have zero TNS and minimum slack 0.000 ns;
  • minimum recovery, removal, and pulse-width slack are respectively +1.056 ns, +0.164 ns, and +0.200 ns.

Hardware acceptance

The accepted core was installed on the DE25-Nano and loaded successfully. The live fabric reported the expected core ID/ABI, healthy status, active HDMI scanout, an advancing frame counter, and zero raster overflow.

The complete f52 scheduler replay then passed on hardware:

  • all 311 epochs completed;
  • all fresh texture CRCs matched and resident reuse checks passed;
  • every epoch reported zero fragment drops;
  • DONE rc=0 (all gates passed);
  • the dump contained 307,200 words, sum32=0xaad0b94d, xor32=0x33138181, 245,155 nonzero pixels, and bounds (0,0)..(511,479).

The retrieved Ch437 board framebuffer has SHA-256 d0047677371a0f6e4e319458926f604a8599c92baaf34d09b8e1fd452e31662b and is byte-identical to the accepted Ch436 framebuffer. This is the intended result: Ch437 changes only the HDMI presentation filter, after framebuffer storage, while preserving the proven renderer and replay output.

Ch437 is accepted. Its closure criteria were:

  • fitter completes within the device's 4,680-LAB capacity;
  • line buffers remain inferred M20Ks and total RAM stays at 344 / 358 blocks;
  • all setup, hold, recovery, removal, and pulse-width classes are nonnegative;
  • the narrow EMIF request-address family remains clean;
  • the 25 MHz presentation domain remains clean.

All five criteria passed before deployment.