Skip to content

LU2-WP03 - 3M Gold Compute

Decision

LU2-WP03 = PASS_WITH_LIMITATIONS and 3M_GOLD_COMPUTE = PASS. The mandatory Gold Compute evidence consists of the existing controlled Silver case (Workload A) and two replays of a materially distinct real FEM case (Workload B). Both replays use the unchanged WP02 configuration freeze.

This is a bounded compute claim, not a general scalability or accuracy claim.

Frozen execution contract

  • Freeze ID: LU2-WP02-FREEZE-bfd1975b012453a3
  • Freeze digest: bfd1975b012453a3b492cc79c968ceeba6ae6951a293e3ce65ddda548d8339a1
  • Runtime: qf-solver-large@sha256:d6a1718001fc36772906d1a9505637bbd0a4b7e1d8ccc9afdbcb6f67b7ff6d0e
  • PETSc 3.25.1, MPICH 5.0.1, 8 MPI ranks, contiguous partition, AIJ, CG, GAMG
  • Solver relative tolerance: 1e-10; WP14 acceptance tolerance: 1e-8
  • Execution source snapshot: 0a6b573485cb39d07b5e179aecd654af41bbc8e7
  • No post-result tuning, formulation change or silent fallback

The predeclared contract is wp03_execution_contract.json.

Workloads

Workload A is the existing WP18 Silver control. It is retained as historical controlled evidence and is not rewritten:

  • 3,000,000 true DOF and 5,821,794 TET4 elements
  • unit-cube structured workload
  • input digest 084a471b1caab628e8558c65b1777692ed53d504baad681bf0985c411a33671b
  • historical execution source 9c0605645fa60ef0d89f3ce98ca361a677f13d1d
  • two existing replays, both PASS

Workload B is a new real FEM workload generated by the command recorded in the contract. It retains the TET4 linear-static homogeneous isotropic route while changing the block aspect ratio and physical dimensions:

  • 99 x 99 x 99 brick cells, six TET4 per cell
  • 1,000,000 nodes, 5,821,794 elements and 3,000,000 true DOF
  • dimensions 2.0 m x 0.75 m x 1.25 m
  • all translations fixed at x=0
  • uniform 1,000,000 N x-direction nodal load on x=length
  • input digest eae54fecd4bf8a6ebebf0363e3103c3defab34f34725b706c6441482d8d8b122

Workload B runs

Run Iterations PC setup [s] KSP solve [s] Total [s] Peak RSS [bytes] Residual Equilibrium Energy error Verdict
1 1046 12.762320016 448.059492200 1751.351795648 2,925,760,512 9.898e-11 2.917e-9 9.624e-13 PASS
2 1046 12.769014678 451.662341780 1802.153195431 2,925,764,608 9.897e-11 2.916e-9 9.656e-13 PASS

Both runs completed without timeout or resource-limited classification. Outputs were finite and contained no NaN or Inf. The numerical acceptance checks, including SPD-compatible CG, passed without changing WP14 tolerances.

Replay

The replay record is wp03_replay_comparison.json. It records identical input, configuration and freeze digests, identical iteration count, and PASS. The maximum recorded numerical relative delta is 6.06e-13; RSS variation is 4,096 bytes. Timing is allowed to vary: assembly/ operator variation is 48.279 s, KSP variation is 3.603 s, and total variation is 50.801 s. No performance conclusion is inferred from those differences.

Workload A/B comparison

The descriptive comparison is wp03_workload_comparison.json. The workloads have the same DOF, element count, route, rank count and frozen configuration, but different geometry and input digests. Workload A recorded 598 iterations and 1,568.651 s total; Workload B recorded 1,046 iterations and 1,751.352 s on replay 1. These values are descriptive only. They are not a speedup or regression benchmark because the physical workloads are distinct.

Measurement boundary

The evidence measures model setup, operator/assembly setup, preconditioner setup, KSP solve, post-processing and total time. Preflight, redistribution, communication and I/O are explicitly NOT_MEASURED by the runner and are not inferred from total time.

The machine-readable index and state are:

Claim boundary and next step

The evidence supports only structured TET4 homogeneous isotropic linear-static FEM on the pinned single-host Docker/PETSc/MPI environment and configuration. It does not claim universal 3M performance, multi-node execution, GPU support, mixed meshes, nonlinear analysis, other element families or restart/checkpoint capability. PETSc and the AIJ memory footprint remain environment-dependent.

LU2-WP03 is complete and LU2-WP04 (5M Bronze) is the next work package.