Empirical Noise Calibration¶
Measuring run-to-run variance of TypeSafe System One (jev-1.13.0) and OpenAI Decisions API (gpt-6-luna) across evaluation presets.
The Measurement Problem¶
Large language model scoring engines can exhibit small score fluctuations across repeated calls for identical inputs. To establish whether a \(\pm 0.10\) threshold safety margin is statistically valid, typesafe-eval measures run-to-run absolute difference:
Empirical Noise Matrix: TypeSafe System One (jev-1.13.0)¶
Measured across 60 total uncached evaluations (3 documents \(\times\) 4 runs across 5 presets, totaling 396 pairwise comparisons against model jev-1.13.0):
| Preset | Comparisons | Mean |Δ| | Median |Δ| | 99th %ile |Δ| | Max |Δ| |
|---|---|---|---|---|---|
quality |
36 | 0.0008 | 0.000 | 0.005 | 0.005 |
safety |
54 | 0.0038 | 0.000 | 0.030 | 0.030 |
tech-spec |
36 | 0.0111 | 0.010 | 0.020 | 0.020 |
design-doc |
180 | 0.0021 | 0.000 | 0.020 | 0.020 |
pr-description |
90 | 0.0086 | 0.010 | 0.040 | 0.040 |
| Overall | 396 | 0.0045 | 0.000 | 0.030 | 0.040 |
Empirical Noise Matrix: OpenAI Decisions API (gpt-6-luna)¶
Measured across 45 total uncached evaluations (3 documents \(\times\) 3 runs across 5 presets, totaling 198 pairwise comparisons against model gpt-6-luna via scripts/measure_noise.py --provider openai):
| Preset | Comparisons | Mean |Δ| | Median |Δ| | 99th %ile |Δ| | Max |Δ| |
|---|---|---|---|---|---|
quality |
18 | 0.0000 | 0.000 | 0.000 | 0.000 |
safety |
27 | 0.0000 | 0.000 | 0.000 | 0.000 |
tech-spec |
18 | 0.0000 | 0.000 | 0.000 | 0.000 |
design-doc |
90 | 0.0000 | 0.000 | 0.000 | 0.000 |
pr-description |
45 | 0.0000 | 0.000 | 0.000 | 0.000 |
| Overall | 198 | 0.0000 | 0.000 | 0.000 | 0.000 |
Multi-Provider Noise Comparison¶
| Dimension | TypeSafe System One (jev-1.13.0) |
OpenAI Decisions API (gpt-6-luna) |
|---|---|---|
| Overall Mean |Δ| | \(0.0045\) | \(0.0000\) |
| 99th %ile |Δ| | \(0.030\) | \(0.000\) |
| Maximum |Δ| | \(0.040\) | \(0.000\) |
| Output Token Cost | Variable | \(0\) tokens |
| Determinism | Statistical (\(\Delta \le 0.04\)) | Strictly Deterministic (\(\Delta = 0.00\)) |
Key Conclusions¶
- Safety Margin Validity:
- For
jev-1.13.0, the 99th percentile noise is \(0.030\) (max \(0.040\)), confirming that the \(\pm 0.10\) threshold margin provides a robust buffer (>2.5× empirical noise). -
For
gpt-6-luna, forward-pass logit extraction is strictly deterministic (\(\Delta = 0.0000\)), eliminating score jitter entirely between identical runs. -
Near-Threshold Annealing:
- Scores landing in \([0.40, 0.60]\) remain soft observations under both models. Capping automated agent rewrite loops at 2 iterations avoids unnecessary over-optimization churn.