30 Jul., 23:14 UTC
FlopScope 0.10.0 released on GitHub.
A regime-aware map of depth-32 estimation: exact identities, deterministic cubature, compute-aware algebra, negative results, and the limits of public-score inference.
The campaign is large, but its counts answer different questions. Collapsing them would turn an audit trail into a misleading headline.
A GitHub release time, a first observed production evaluation, and an official announcement are different events. The analysis retains each rather than forcing them into one change point.
FlopScope 0.10.0 released on GitHub.
WhestBench 0.14.0 released on GitHub.
First final-runtime evaluation observed in this reconstruction.
AIcrowd publicly announced the new runtime, cost-model fixes, and re-evaluation [official update].
The interactive plot contains all 238 currently graded rows with exposed positive score and compute. Filter by analytical campaign; keyboard and screen-reader users can use the synchronized top-record table.
| Submission | Score | Compute | Campaign |
|---|
Analytic closures, learned correctors, covariance compression, multifidelity controls, conditional integration, MLMC, QMC variants, and native-kernel paths were tested against explicit gates.
One long Sobol stream with exact antipodal pairing beat eight shifted replicates by 7.04% in a 500-network, three-seed screen. Stronger LMS scrambling regressed.
Row counts, rotations, frequency packs, pilot sizes, and thresholds were treated as configurations inside one broad design family, not independent theories.
The late frontier held raw MSE near 1.92e-7 while Walsh-Hadamard, antipodal, and Strassen-Winograd rewrites lowered metered compute.
The final estimator combines a structured directional design with exact metered algebra. It is not “all exact”: the pilot masks, live-column pruning, finite cubature, pilot-dead zero fill, and pilot-row reuse are approximation choices.
For X=RU and a bias-free positively homogeneous ReLU network f, f(RU)=R f(U). Estimate the sphere expectation and multiply by E[R].
Use a weighted Kerdock/MUB carrier. A numerical certificate checks its quantized low-order moment identities; this does not prove exactness for the deep nonlinear integrand.
Load the shipped 1,152-row pilot (512 rotated axis rows and a 640-row code supplement). Keep a coordinate when it fires on more than two pilot rows.
Factor each Kerdock coset into sign, Walsh-Hadamard, and diagonal operations; replace dense layer-0 multiplication with one 256-cubed product plus FWHTs.
Compute one side of each +/- pair and recover the mirror using ReLU(-y)=ReLU(y)-y, with an additional transform for the layer-1 mirror.
Propagate 64,896 main code trajectories through layers 2-30 after removing pilot-inactive columns; combine them later with the 640 pilot code rows and 512 axis rows to recover the weighted 66,048-point estimate.
For each live matrix shape, compare dense, padding, peel-k, peel-n, and peel-both plans across Strassen-Winograd depths; all fringe decompositions are exact.
Combine an always-on linear mean, a row-wise ReLU path for kinked units, zero-fill pilot-dead units, and reuse the pilot's 512 axis rows plus 640 code rows in the weighted final mean.
implementation asset: 512 rotated axis rows followed by a 640-row code supplement. Its live-mask condition is count > 2. Later loose source used a different pilot path and is not evidence for the submitted artifact.Five byte-identical repeats span 8.69896e-8 to 8.71106e-8, a 0.139% range relative to their mean. That bounds repeat noise for this artifact under the final captured public regime; it does not bound fresh-network uncertainty.
Each row is phrased as question → controlled observation → boundary. A boundary applies only to the measured implementation and evaluation surface.
| Question | Controlled observation | Reusable boundary |
|---|---|---|
| Can Gaussian moment propagation replace sampling at depth 32? | Final-layer MSE was 7.5133e-5 at width 256, about 411 times the contemporaneous 1.83e-7 reference; error accumulated with depth. | Not competitive alone in this regime; useful as a mechanism and scaling diagnostic. |
| Does covariance spectral concentration justify truncation? | Effective rank fell from about 164 to 2.8, but ranks 8-64 left relative error nearly flat near 3.4e-2. | Low effective rank did not imply a safe low-rank covariance approximation. |
| Can the omitted diagonal residual image be sketched? | A W^T diag(d) W sketch improved relative L2 error from 1.29e-2 to 6.9e-3 at r=64, k=16. | A positive component result, but the estimated full-score effect was only 1.7-3%. |
| Do stronger digital scrambles automatically improve QMC? | Two LMS-scrambled banks regressed by 3.25% and 5.88% in the controlled local screen. | Scramble quality was integrand-specific; sophistication alone was not a promotion criterion. |
| Can active-subspace stratification help? | The q2 arm improved 10.09% over 32 networks x 3 observations, but its clustered interval crossed zero. | Promising signal, insufficient generalization evidence. |
| Can exact conditional integration remove variance cheaply? | The best ray direction removed 20.64% of variance but required roughly 8,000-10,000 regions and 6-7 seconds per ray. | Mathematically effective, computationally infeasible in the measured implementation. |
| Do structural or learned controls produce a large gain? | Most measured gains were below 0.1-3%; honest cross-fit variants sometimes regressed. | The tested control families did not clear their predeclared cost-normalized gates. |
| Does exact MLMC telescoping guarantee favorable economics? | The pathwise identity held, but projected raw MSE was 5.26e-6, more than 11 times the target. | An exact identity can still have an unfavorable level-variance budget. |
| Can lower precision or VNNI supply the missing factor? | The tested capability marker did not fire: the evaluated environment did not expose AVX-512 VNNI to that submission. The tested block-int8 path took 7.1686 s versus 0.18745 s for FP32 (38.24x slower) and 24.98x the separate 0.287 s residual-time target. | This specific VNNI path was not exposed to the tested submission, and this block-int8 implementation failed its measured throughput target despite acceptable numerical fidelity. |
| Does higher cumulant order retain its advantage with depth? | Corrected sweeps had k3 beat k2 at every tested depth, but the advantage decayed from 35x at depth 8 to 1.8x at depth 32; k4 ran out of memory at depth 24 and above on 48 GB. | Higher order still helped, but depth and memory bound the approach. |
After hundreds of submissions, the public suite is selection data. Public-only resampling can characterize instability on those same networks; it cannot manufacture fresh MLPs or seeds.
These buttons reconstruct embedded byte-for-byte copies, so the page remains useful offline. The adjacent package also contains chart-level CSV/JSON and SVG/PNG assets.
The public timeline keeps method and outcome rows while excluding internal provenance handles. The 332-row grader appendix keeps current status, score fields, failure messages, runtime labels, method labels, and public submission URLs. Source packages, authentication material, and other participants' code are not embedded in this public page.
OpenAI Codex and Anthropic Claude assisted with orchestration, code generation, evidence indexing, data analysis, figure production, and drafting; GLM supplied an additional proofreading and review pass. Human direction determined questions, promotion/kill decisions, candidate selection, and publication scope. Model agreement was not treated as validation: outputs were challenged by other models and then checked against preserved official grader records and contemporaneous participant records. Reconstructed labels remain distinguished from directly checked submitted configurations.
AIcrowd participant pluto's “keep going” account inspired the persistence side of this workflow. My variation added a friendly Simpsons-style pile-on: Codex, Claude, and GLM were asked to challenge one another's claims, with the recorded results deciding disagreements.