4e8c3a16b8
The per-triangle gradient engine already time-shared ONE divider across all GRAD_STEPS attributes, but computed every grad_load_num[0:GRAD_STEPS-1] numerator IN PARALLEL — ~44 wide multiplies (~100 physical DSP) for once-per-triangle setup consumed one-at-a-time. Pure redundant hardware; the design was DSP-maxed (187/188, 99%) so nothing new could fit (fog needed 191/188). Replace the parallel bank + the grad_num_q[] pre-latch array with grad_num_step: computes ONLY the current grad_step's numerator from ONE mux-selected pair of signed multipliers (attribute triple by grad_step>>1, axis by grad_step[0]; shared da1/da2, two shared products, signed subtract, <<<20). grad_word_q/grad_slot are held stable the whole solve, so it is bit-identical to the old grad_num_q[grad_step]. Removed grad_num_dadx/dady (inlined once). FSM sequencing and throughput unchanged. Width note: da1/da2 are 33-bit (products 50-bit), NOT operand-width 32-bit — the original (a1-a0) lived in a signed-64-bit expression context and never wrapped; full-32-bit Z with |a1-a0|>2^31 needs the wider intermediate. tb_gs_grad_num_equiv (extreme signed corners + 200k random = 494770 checks, 0 errors) caught a 32-bit first cut that f52's real data never exercised. Resource (26.1 Seed-3 fit): DSP needed 168->83 / final placement 187->119, i.e. 99% -> 44%, ~85 blocks reclaimed (Codex gate >=70 met). ALM 40458->38784 (86->83%). RAM 322/358 unchanged. Timing CLEAN: setup +0.077, all classes >=0, 0 violated. Verification: tb_gs_grad_num_equiv 0/494770; f52 replay BYTE-IDENTICAL golden d0047677 (drops=0, occupancy unchanged); gradient/perspective/texture regressions (tri_interp, grad_divider, persp_uv, zbuffer, fog_persp, textured_triangle, triangle/perspective/combined/gouraud demos) all PASS. Byte-identical => the screen is unchanged; this is the resource unlock for fog + coverage. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>