Insights DCLASS: Pipedrive schema qualification

Overall frozen acceptance: **FAIL**. Target real correctness changes by 0.00 questions per four-question repeat (0.00 percentage points), with paired family-cluster 95% interval [0.00, 0.00] questions. The control observed repeat range is 0.00 questions; the candidate range is 0.00. 0 of three candidate repeats exceed the control mean.

Target real silent-wrong answers change from 0.67 to 2.33 per repeat: difference 1.67 [0.67, 2.67] questions, against a control range of 1.00. The whole-sample and complementary-set safety gates are aggregate tests; their PASS labels do not establish target safety or outweigh failed real-question correctness. The target tables below must be read alongside those aggregate gates.

The production change preserves the Pipedrive schema identifier at the shared generation-context loader. It changes no generator, verifier, repair policy, oracle, source allowlist or business predicate. The observed-defect target was frozen from D3 records: four real first-turn questions across P02/P03/P05/X01 and sixteen fresh cases across P01–P05. Broader family populations are reported separately.

This is a development-informed comparison against historical fresh controls and mixed-date real controls. The target was selected from historical failures, so regression to the mean is possible. The family intervals condition on the observed repeats and do not model provider or run-level uncertainty. Every latency below was measured on a shared host; historical and new timing conditions differ.

Per-repeat outcomes

First-turn errors use supported first turns. Total errors include task and infrastructure failures on every turn. Refused combines unverified and justified refusals; the exhaustive outcome and guard partitions follow. Precision excludes comparison-unverified answers. Counts, means and ranges describe completed evaluator attempts.

Dataset / setArm / repeatAttemptsFirst turns / supportedFirst-turn errorsCorrectWrongRefusedComparison unverifiedTotal errorsPrecisionShared-host p95 (s)
fresh / wholecontrol-fresh-r1122116 / 116301293822323.33%29.43
fresh / wholecontrol-fresh-r2122116 / 1162732741232810.00%34.54
fresh / wholecontrol-fresh-r3122116 / 116273343819288.11%27.29
fresh / wholecandidate-fresh-r1122116 / 116192254926207.41%24.98
fresh / wholecandidate-fresh-r2122116 / 1161932744291910.00%26.52
fresh / wholecandidate-fresh-r3122116 / 116163294330179.38%32.05
fresh / targetcontrol-fresh-r11616 / 1614001114n/a32.48
fresh / targetcontrol-fresh-r21616 / 1613100213100.00%45.04
fresh / targetcontrol-fresh-r31616 / 16130111130.00%42.67
fresh / targetcandidate-fresh-r11616 / 1622336240.00%24.98
fresh / targetcandidate-fresh-r21616 / 1621139250.00%36.91
fresh / targetcandidate-fresh-r31616 / 1622327240.00%41.24
fresh / restcontrol-fresh-r1106100 / 100161293721183.33%29.43
fresh / restcontrol-fresh-r2106100 / 100142274121156.90%32.55
fresh / restcontrol-fresh-r3106100 / 100143333718158.33%27.29
fresh / restcandidate-fresh-r1106100 / 100170224620180.00%25.10
fresh / restcandidate-fresh-r2106100 / 100172264120177.14%25.28
fresh / restcandidate-fresh-r3106100 / 100141264123153.70%32.05
real / wholecontrol-real-r110191 / 86101223625174.35%31.20
real / wholecontrol-real-r210191 / 86102263027167.14%27.82
real / wholecontrol-real-r310191 / 86102332426165.71%26.72
real / wholecandidate-real-r110191 / 8672282730146.67%24.25
real / wholecandidate-real-r210191 / 86643325271210.81%25.86
real / wholecandidate-real-r310191 / 86732731281210.00%25.06
real / targetcontrol-real-r144 / 4400004n/a21.13
real / targetcontrol-real-r244 / 42010120.00%21.94
real / targetcontrol-real-r344 / 43010030.00%29.47
real / targetcandidate-real-r144 / 41020110.00%24.25
real / targetcandidate-real-r244 / 40030100.00%29.33
real / targetcandidate-real-r344 / 41020110.00%27.95
real / restcontrol-real-r19787 / 8261223625134.35%35.63
real / restcontrol-real-r29787 / 8282253026147.41%29.77
real / restcontrol-real-r39787 / 8272322426135.88%26.72
real / restcandidate-real-r19787 / 8262262729137.14%25.58
real / restcandidate-real-r29787 / 82643025261211.76%24.08
real / restcandidate-real-r39787 / 82632531271110.71%25.06

Repeat means and ranges

Cells are mean [minimum, maximum] over defined repeats; precision also states its defined-repeat n. Count summaries use all three repeats. Target-control precision is undefined in repeat 1 for both datasets, so those two means use n=2; undefined precision is not replaced by zero. The control max−min range is the acceptance noise floor, an observed range rather than a confidence bound. Candidate ranges are also shown in the noise table.

Dataset / setArmFirst-turn errorsCorrectWrongRefusedComparison unverifiedTotal errorsPrecisionShared-host p95 (s)
fresh / wholecontrol28.00 [27.00, 30.00]2.33 [1.00, 3.00]30.00 [27.00, 34.00]39.00 [38.00, 41.00]21.33 [19.00, 23.00]29.33 [28.00, 32.00]7.15% [3.33%, 10.00%] (n=3)30.42 [27.29, 34.54]
fresh / wholecandidate18.00 [16.00, 19.00]2.67 [2.00, 3.00]27.00 [25.00, 29.00]45.33 [43.00, 49.00]28.33 [26.00, 30.00]18.67 [17.00, 20.00]8.93% [7.41%, 10.00%] (n=3)27.85 [24.98, 32.05]
fresh / targetcontrol13.33 [13.00, 14.00]0.33 [0.00, 1.00]0.33 [0.00, 1.00]0.67 [0.00, 1.00]1.33 [1.00, 2.00]13.33 [13.00, 14.00]50.00% [0.00%, 100.00%] (n=2)40.06 [32.48, 45.04]
fresh / targetcandidate2.00 [2.00, 2.00]1.67 [1.00, 2.00]2.33 [1.00, 3.00]2.67 [2.00, 3.00]7.33 [6.00, 9.00]2.00 [2.00, 2.00]43.33% [40.00%, 50.00%] (n=3)34.38 [24.98, 41.24]
fresh / restcontrol14.67 [14.00, 16.00]2.00 [1.00, 3.00]29.67 [27.00, 33.00]38.33 [37.00, 41.00]20.00 [18.00, 21.00]16.00 [15.00, 18.00]6.19% [3.33%, 8.33%] (n=3)29.76 [27.29, 32.55]
fresh / restcandidate16.00 [14.00, 17.00]1.00 [0.00, 2.00]24.67 [22.00, 26.00]42.67 [41.00, 46.00]21.00 [20.00, 23.00]16.67 [15.00, 18.00]3.62% [0.00%, 7.14%] (n=3)27.48 [25.10, 32.05]
real / wholecontrol10.00 [10.00, 10.00]1.67 [1.00, 2.00]27.00 [22.00, 33.00]30.00 [24.00, 36.00]26.00 [25.00, 27.00]16.33 [16.00, 17.00]5.73% [4.35%, 7.14%] (n=3)28.58 [26.72, 31.20]
real / wholecandidate6.67 [6.00, 7.00]3.00 [2.00, 4.00]29.33 [27.00, 33.00]27.67 [25.00, 31.00]28.33 [27.00, 30.00]12.67 [12.00, 14.00]9.16% [6.67%, 10.81%] (n=3)25.06 [24.25, 25.86]
real / targetcontrol3.00 [2.00, 4.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]3.00 [2.00, 4.00]0.00% [0.00%, 0.00%] (n=2)24.18 [21.13, 29.47]
real / targetcandidate0.67 [0.00, 1.00]0.00 [0.00, 0.00]2.33 [2.00, 3.00]0.00 [0.00, 0.00]1.00 [1.00, 1.00]0.67 [0.00, 1.00]0.00% [0.00%, 0.00%] (n=3)27.18 [24.25, 29.33]
real / restcontrol7.00 [6.00, 8.00]1.67 [1.00, 2.00]26.33 [22.00, 32.00]30.00 [24.00, 36.00]25.67 [25.00, 26.00]13.33 [13.00, 14.00]5.88% [4.35%, 7.41%] (n=3)30.71 [26.72, 35.63]
real / restcandidate6.00 [6.00, 6.00]3.00 [2.00, 4.00]27.00 [25.00, 30.00]27.67 [25.00, 31.00]27.33 [26.00, 29.00]12.00 [11.00, 13.00]9.87% [7.14%, 11.76%] (n=3)24.91 [24.08, 25.58]

Outcome and refusal partitions

Each outcome partition sums to that repeat’s attempt denominator; each guard partition sums to its refusal count. No error, blocked turn, clarification or unverified result is silently counted as a refusal or a correct answer.

Dataset / setRepeatExhaustive outcomesRefusals by guard
fresh / wholecontrol-fresh-r1{"comparison_unverified": 22, "correct_completed": 1, "refusal_unverified": 38, "silent_wrong": 29, "task_error": 32}{"CalendarSemanticsError": 33, "NativeDealUnitError": 4, "UNKNOWN-GOLD-COLUMN": 1}
fresh / wholecontrol-fresh-r2{"comparison_unverified": 23, "correct_completed": 3, "refusal_unverified": 41, "silent_wrong": 27, "task_error": 28}{"AdditiveGrainError": 1, "CalendarSemanticsError": 39, "NativeDealUnitError": 1}
fresh / wholecontrol-fresh-r3{"comparison_unverified": 19, "correct_completed": 3, "refusal_unverified": 38, "silent_wrong": 34, "task_error": 28}{"CalendarSemanticsError": 36, "NativeDealUnitError": 2}
fresh / wholecandidate-fresh-r1{"comparison_unverified": 26, "correct_completed": 2, "refusal_unverified": 49, "silent_wrong": 25, "task_error": 20}{"CalendarSemanticsError": 43, "LossReasonSourceError": 1, "NativeDealUnitError": 5}
fresh / wholecandidate-fresh-r2{"comparison_unverified": 29, "correct_completed": 3, "refusal_unverified": 44, "silent_wrong": 27, "task_error": 19}{"AdditiveGrainError": 1, "CalendarSemanticsError": 38, "NativeDealUnitError": 5}
fresh / wholecandidate-fresh-r3{"comparison_unverified": 30, "correct_completed": 3, "refusal_unverified": 43, "silent_wrong": 29, "task_error": 17}{"AdditiveGrainError": 1, "CalendarSemanticsError": 38, "FanOutSumError": 1, "NativeDealUnitError": 3}
fresh / targetcontrol-fresh-r1{"comparison_unverified": 1, "refusal_unverified": 1, "task_error": 14}{"NativeDealUnitError": 1}
fresh / targetcontrol-fresh-r2{"comparison_unverified": 2, "correct_completed": 1, "task_error": 13}{}
fresh / targetcontrol-fresh-r3{"comparison_unverified": 1, "refusal_unverified": 1, "silent_wrong": 1, "task_error": 13}{"NativeDealUnitError": 1}
fresh / targetcandidate-fresh-r1{"comparison_unverified": 6, "correct_completed": 2, "refusal_unverified": 3, "silent_wrong": 3, "task_error": 2}{"NativeDealUnitError": 3}
fresh / targetcandidate-fresh-r2{"comparison_unverified": 9, "correct_completed": 1, "refusal_unverified": 3, "silent_wrong": 1, "task_error": 2}{"NativeDealUnitError": 3}
fresh / targetcandidate-fresh-r3{"comparison_unverified": 7, "correct_completed": 2, "refusal_unverified": 2, "silent_wrong": 3, "task_error": 2}{"NativeDealUnitError": 2}
fresh / restcontrol-fresh-r1{"comparison_unverified": 21, "correct_completed": 1, "refusal_unverified": 37, "silent_wrong": 29, "task_error": 18}{"CalendarSemanticsError": 33, "NativeDealUnitError": 3, "UNKNOWN-GOLD-COLUMN": 1}
fresh / restcontrol-fresh-r2{"comparison_unverified": 21, "correct_completed": 2, "refusal_unverified": 41, "silent_wrong": 27, "task_error": 15}{"AdditiveGrainError": 1, "CalendarSemanticsError": 39, "NativeDealUnitError": 1}
fresh / restcontrol-fresh-r3{"comparison_unverified": 18, "correct_completed": 3, "refusal_unverified": 37, "silent_wrong": 33, "task_error": 15}{"CalendarSemanticsError": 36, "NativeDealUnitError": 1}
fresh / restcandidate-fresh-r1{"comparison_unverified": 20, "refusal_unverified": 46, "silent_wrong": 22, "task_error": 18}{"CalendarSemanticsError": 43, "LossReasonSourceError": 1, "NativeDealUnitError": 2}
fresh / restcandidate-fresh-r2{"comparison_unverified": 20, "correct_completed": 2, "refusal_unverified": 41, "silent_wrong": 26, "task_error": 17}{"AdditiveGrainError": 1, "CalendarSemanticsError": 38, "NativeDealUnitError": 2}
fresh / restcandidate-fresh-r3{"comparison_unverified": 23, "correct_completed": 1, "refusal_unverified": 41, "silent_wrong": 26, "task_error": 15}{"AdditiveGrainError": 1, "CalendarSemanticsError": 38, "FanOutSumError": 1, "NativeDealUnitError": 1}
real / wholecontrol-real-r1{"comparison_unverified": 25, "correct_completed": 1, "refusal_unverified": 36, "silent_wrong": 22, "task_error": 17}{"-": 1, "CalendarSemanticsError": 32, "NativeDealUnitError": 2, "WindowBindingError": 1}
real / wholecontrol-real-r2{"comparison_unverified": 27, "correct_completed": 2, "refusal_unverified": 30, "silent_wrong": 26, "task_error": 16}{"-": 1, "CalendarSemanticsError": 28, "NativeDealUnitError": 1}
real / wholecontrol-real-r3{"comparison_unverified": 26, "correct_completed": 2, "refusal_unverified": 24, "silent_wrong": 33, "task_error": 16}{"-": 1, "CalendarSemanticsError": 21, "NativeDealUnitError": 2}
real / wholecandidate-real-r1{"comparison_unverified": 30, "correct_completed": 2, "refusal_unverified": 27, "silent_wrong": 28, "task_error": 14}{"-": 1, "CalendarSemanticsError": 25, "NativeDealUnitError": 1}
real / wholecandidate-real-r2{"comparison_unverified": 27, "correct_completed": 4, "refusal_unverified": 25, "silent_wrong": 33, "task_error": 12}{"-": 1, "CalendarSemanticsError": 23, "NativeDealUnitError": 1}
real / wholecandidate-real-r3{"comparison_unverified": 28, "correct_completed": 3, "refusal_unverified": 31, "silent_wrong": 27, "task_error": 12}{"-": 1, "AdditiveGrainError": 1, "CalendarSemanticsError": 25, "FactDateScopeError": 1, "NativeDealUnitError": 2, "WindowBindingError": 1}
real / targetcontrol-real-r1{"task_error": 4}{}
real / targetcontrol-real-r2{"comparison_unverified": 1, "silent_wrong": 1, "task_error": 2}{}
real / targetcontrol-real-r3{"silent_wrong": 1, "task_error": 3}{}
real / targetcandidate-real-r1{"comparison_unverified": 1, "silent_wrong": 2, "task_error": 1}{}
real / targetcandidate-real-r2{"comparison_unverified": 1, "silent_wrong": 3}{}
real / targetcandidate-real-r3{"comparison_unverified": 1, "silent_wrong": 2, "task_error": 1}{}
real / restcontrol-real-r1{"comparison_unverified": 25, "correct_completed": 1, "refusal_unverified": 36, "silent_wrong": 22, "task_error": 13}{"-": 1, "CalendarSemanticsError": 32, "NativeDealUnitError": 2, "WindowBindingError": 1}
real / restcontrol-real-r2{"comparison_unverified": 26, "correct_completed": 2, "refusal_unverified": 30, "silent_wrong": 25, "task_error": 14}{"-": 1, "CalendarSemanticsError": 28, "NativeDealUnitError": 1}
real / restcontrol-real-r3{"comparison_unverified": 26, "correct_completed": 2, "refusal_unverified": 24, "silent_wrong": 32, "task_error": 13}{"-": 1, "CalendarSemanticsError": 21, "NativeDealUnitError": 2}
real / restcandidate-real-r1{"comparison_unverified": 29, "correct_completed": 2, "refusal_unverified": 27, "silent_wrong": 26, "task_error": 13}{"-": 1, "CalendarSemanticsError": 25, "NativeDealUnitError": 1}
real / restcandidate-real-r2{"comparison_unverified": 26, "correct_completed": 4, "refusal_unverified": 25, "silent_wrong": 30, "task_error": 12}{"-": 1, "CalendarSemanticsError": 23, "NativeDealUnitError": 1}
real / restcandidate-real-r3{"comparison_unverified": 27, "correct_completed": 3, "refusal_unverified": 31, "silent_wrong": 25, "task_error": 11}{"-": 1, "AdditiveGrainError": 1, "CalendarSemanticsError": 25, "FactDateScopeError": 1, "NativeDealUnitError": 2, "WindowBindingError": 1}

Paired intervals and observed noise

Candidate minus control; count differences are normalized to one repeat of the named set. 10,000 whole-family resamples, seed 2026090611, nearest-rank 2.5th and 97.5th percentiles. Each arm contributes its repeat-mean family counts; repeat indices are not causal pairs. Precision and p95 are pooled within the bootstrap and may differ from differences of repeat means.

Dataset / setMetricDifference [95% interval]Control spreadCandidate spreadUndefined draws
fresh / wholefirst_turn_execution_errors-10.00 [-22.19, 0.35]3.003.000
fresh / wholecorrect_completed0.33 [-1.68, 3.02]2.001.000
fresh / wholesilent_wrong-3.00 [-8.07, 2.02]7.004.000
fresh / wholerefusals6.33 [-0.68, 12.44]3.006.000
fresh / wholerefusals_by_guard6.33 [-0.68, 12.44]3.006.000
fresh / wholerefusals_by_verifier0.00 [0.00, 0.00]0.000.000
fresh / wholecomparison_unverified7.00 [0.32, 14.82]4.004.000
fresh / wholefull_conversation_successes0.00 [0.00, 0.00]0.000.00413
fresh / wholeanswer_precision0.02 [-0.05, 0.09]0.070.030
fresh / wholep95_seconds-4.32 [-7.84, 1.83]7.257.080
fresh / targetfirst_turn_execution_errors-11.33 [-13.93, -7.67]1.000.000
fresh / targetcorrect_completed1.33 [0.00, 2.82]1.001.000
fresh / targetsilent_wrong2.00 [0.63, 3.26]1.002.000
fresh / targetrefusals2.00 [0.33, 3.91]1.001.000
fresh / targetrefusals_by_guard2.00 [0.33, 3.91]1.001.000
fresh / targetrefusals_by_verifier0.00 [0.00, 0.00]0.000.000
fresh / targetcomparison_unverified6.00 [3.05, 8.33]1.003.000
fresh / targetfull_conversation_successesn/a0.000.0010000
fresh / targetanswer_precision-0.08 [-0.67, 0.38]1.000.10737
fresh / targetp95_seconds-9.02 [-21.17, 10.94]12.5616.270
fresh / restfirst_turn_execution_errors1.33 [-2.24, 5.96]2.003.000
fresh / restcorrect_completed-1.00 [-2.12, 0.00]2.002.000
fresh / restsilent_wrong-5.00 [-9.35, -0.75]6.004.000
fresh / restrefusals4.33 [-2.21, 10.10]4.005.000
fresh / restrefusals_by_guard4.33 [-2.21, 10.10]4.005.000
fresh / restrefusals_by_verifier0.00 [0.00, 0.00]0.000.000
fresh / restcomparison_unverified1.00 [-2.94, 5.20]3.003.000
fresh / restfull_conversation_successes0.00 [0.00, 0.00]0.000.00420
fresh / restanswer_precision-0.02 [-0.07, 0.01]0.050.070
fresh / restp95_seconds-2.14 [-6.50, 1.93]5.266.950
real / wholefirst_turn_execution_errors-3.33 [-7.64, 0.00]0.001.000
real / wholecorrect_completed1.33 [-0.33, 3.21]1.002.000
real / wholesilent_wrong2.33 [-2.00, 6.99]11.006.000
real / wholerefusals-2.33 [-6.45, 1.72]12.006.000
real / wholerefusals_by_guard-2.33 [-6.45, 1.72]12.006.000
real / wholerefusals_by_verifier0.00 [0.00, 0.00]0.000.000
real / wholecomparison_unverified2.33 [-2.83, 7.80]2.003.000
real / wholefull_conversation_successes0.00 [0.00, 0.00]0.000.0046
real / wholeanswer_precision0.03 [-0.01, 0.08]0.030.040
real / wholep95_seconds-4.34 [-6.81, -0.28]4.481.610
real / targetfirst_turn_execution_errors-2.33 [-2.67, -1.67]2.001.000
real / targetcorrect_completed0.00 [0.00, 0.00]0.000.000
real / targetsilent_wrong1.67 [0.67, 2.67]1.001.000
real / targetrefusals0.00 [0.00, 0.00]0.000.000
real / targetrefusals_by_guard0.00 [0.00, 0.00]0.000.000
real / targetrefusals_by_verifier0.00 [0.00, 0.00]0.000.000
real / targetcomparison_unverified0.67 [0.00, 2.00]1.000.000
real / targetfull_conversation_successesn/a0.000.0010000
real / targetanswer_precision0.00 [0.00, 0.00]0.000.00617
real / targetp95_seconds-0.13 [-1.52, 7.39]8.335.090
real / restfirst_turn_execution_errors-1.00 [-3.85, 1.30]2.000.000
real / restcorrect_completed1.33 [-0.32, 3.20]1.002.000
real / restsilent_wrong0.67 [-3.23, 4.67]10.005.000
real / restrefusals-2.33 [-6.39, 1.70]12.006.000
real / restrefusals_by_guard-2.33 [-6.39, 1.70]12.006.000
real / restrefusals_by_verifier0.00 [0.00, 0.00]0.000.000
real / restcomparison_unverified1.67 [-3.39, 6.81]1.003.000
real / restfull_conversation_successes0.00 [0.00, 0.00]0.000.0053
real / restanswer_precision0.04 [-0.01, 0.10]0.030.050
real / restp95_seconds-5.44 [-7.52, -1.07]8.911.500

Frozen pass targets

TargetVerdictEvidence
target_real_correctnessFAIL{"at_least_two_repeats_above_control_mean": false, "candidate_repeats_above_control_mean": 0, "candidate_spread": 0.0, "ci95": [0.0, 0.0], "ci95_percentage_points": [0.0, 0.0], "control_spread": 0.0, "gain": 0.0, "gain_exceeds_control_spread": false, "gain_percentage_points": 0.0, "paired_lower_bound_positive": false}
whole_fresh_safetyPASS{"components": {"correct_completed": {"adverse_difference_exceeds_spread": false, "ci95": [-1.6804407713498635, 3.0247933884297526], "control_spread": 2.0, "interval_wholly_adverse": false, "mean_delta": 0.33333333333333304, "pass": true}, "silent_wrong": {"adverse_difference_exceeds_spread": false, "ci95": [-8.066115702479344, 2.0165289256198307], "control_spread": 7.0, "interval_wholly_adverse": false, "mean_delta": -3.0, "pass": true}}}
rest_fresh_safetyPASS{"components": {"correct_completed": {"adverse_difference_exceeds_spread": false, "ci95": [-2.1199999999999997, 0.0], "control_spread": 2.0, "interval_wholly_adverse": false, "mean_delta": -1.0, "pass": true}, "silent_wrong": {"adverse_difference_exceeds_spread": false, "ci95": [-9.35294117647059, -0.7517730496453972], "control_spread": 6.0, "interval_wholly_adverse": false, "mean_delta": -5.0, "pass": true}}}
whole_real_safetyPASS{"components": {"correct_completed": {"adverse_difference_exceeds_spread": false, "ci95": [-0.3300653594771239, 3.2063492063492056], "control_spread": 1.0, "interval_wholly_adverse": false, "mean_delta": 1.3333333333333333, "pass": true}, "silent_wrong": {"adverse_difference_exceeds_spread": false, "ci95": [-2.003968253968246, 6.987421383647803], "control_spread": 11.0, "interval_wholly_adverse": false, "mean_delta": 2.333333333333332, "pass": true}}}
rest_real_safetyPASS{"components": {"correct_completed": {"adverse_difference_exceeds_spread": false, "ci95": [-0.3201320132013201, 3.2013201320132008], "control_spread": 1.0, "interval_wholly_adverse": false, "mean_delta": 1.3333333333333333, "pass": true}, "silent_wrong": {"adverse_difference_exceeds_spread": false, "ci95": [-3.2333333333333343, 4.670370370370364], "control_spread": 10.0, "interval_wholly_adverse": false, "mean_delta": 0.6666666666666679, "pass": true}}}
curated_invariantsPASS{"errors": 0, "executed_passed": 636, "failed": 0, "missing_inherited_files": 19, "skipped": 88}

The target real correct-answer gain is inside or at the observed control noise floor. The frozen target correctness requirement therefore fails, regardless of whether the schema bug is removed.

A safety PASS means the frozen adverse-difference tests did not trigger; it is not proof of equivalence. Whole fresh and whole real apply the same safety checks as both complements. The target real conjunction also requires two improved candidate repeats. Its four family clusters provide limited resolution: gains concentrated in only one or two families may not clear the interval.

Affected-family before and after

The first-turn metric counts distinct first-turn cases and excludes derivative turns. All-attempt completion is shown separately. Cells are means [minimum, maximum], then a fixed denominator. A family absent from a set has denominator zero and an undefined rate.

Dataset / membershipFamilyArmCorrect first-turn questions / first turnsCorrect all attempts / attemptsWrongRefusedTotal errors
fresh / targetP01control0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 20.33 [0.00, 1.00]0.67 [0.00, 1.00]1.00 [0.00, 2.00]
fresh / targetP01candidate0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 20.67 [0.00, 1.00]1.33 [1.00, 2.00]0.00 [0.00, 0.00]
fresh / broader_familyP01control0.67 [0.00, 1.00] / 40.67 [0.00, 1.00] / 40.33 [0.00, 1.00]1.67 [1.00, 3.00]1.33 [1.00, 2.00]
fresh / broader_familyP01candidate0.67 [0.00, 1.00] / 40.67 [0.00, 1.00] / 40.67 [0.00, 1.00]2.67 [2.00, 3.00]0.00 [0.00, 0.00]
fresh / targetP02control0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.00 [0.00, 0.00]0.00 [0.00, 0.00]2.67 [2.00, 3.00]
fresh / targetP02candidate0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.33 [0.00, 1.00]0.33 [0.00, 1.00]0.00 [0.00, 0.00]
fresh / broader_familyP02control0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.00 [0.00, 0.00]0.00 [0.00, 0.00]2.67 [2.00, 3.00]
fresh / broader_familyP02candidate0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.33 [0.00, 1.00]0.33 [0.00, 1.00]0.00 [0.00, 0.00]
fresh / targetP03control0.33 [0.00, 1.00] / 40.33 [0.00, 1.00] / 40.00 [0.00, 0.00]0.00 [0.00, 0.00]3.67 [3.00, 4.00]
fresh / targetP03candidate0.67 [0.00, 1.00] / 40.67 [0.00, 1.00] / 40.00 [0.00, 0.00]1.00 [1.00, 1.00]0.67 [0.00, 1.00]
fresh / broader_familyP03control0.33 [0.00, 1.00] / 40.33 [0.00, 1.00] / 40.00 [0.00, 0.00]0.00 [0.00, 0.00]3.67 [3.00, 4.00]
fresh / broader_familyP03candidate0.67 [0.00, 1.00] / 40.67 [0.00, 1.00] / 40.00 [0.00, 0.00]1.00 [1.00, 1.00]0.67 [0.00, 1.00]
fresh / targetP04control0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.00 [0.00, 0.00]0.00 [0.00, 0.00]2.33 [2.00, 3.00]
fresh / targetP04candidate0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.33 [0.00, 1.00]0.00 [0.00, 0.00]1.33 [1.00, 2.00]
fresh / broader_familyP04control0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.00 [0.00, 0.00]0.00 [0.00, 0.00]2.33 [2.00, 3.00]
fresh / broader_familyP04candidate0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 30.33 [0.00, 1.00]0.00 [0.00, 0.00]1.33 [1.00, 2.00]
fresh / targetP05control0.00 [0.00, 0.00] / 40.00 [0.00, 0.00] / 40.00 [0.00, 0.00]0.00 [0.00, 0.00]3.67 [3.00, 4.00]
fresh / targetP05candidate1.00 [1.00, 1.00] / 41.00 [1.00, 1.00] / 41.00 [1.00, 1.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
fresh / broader_familyP05control0.00 [0.00, 0.00] / 40.00 [0.00, 0.00] / 40.00 [0.00, 0.00]0.00 [0.00, 0.00]3.67 [3.00, 4.00]
fresh / broader_familyP05candidate1.00 [1.00, 1.00] / 41.00 [1.00, 1.00] / 41.00 [1.00, 1.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
fresh / targetX01control0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
fresh / targetX01candidate0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
fresh / broader_familyX01control0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 31.67 [1.00, 2.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
fresh / broader_familyX01candidate0.00 [0.00, 0.00] / 30.00 [0.00, 0.00] / 31.33 [1.00, 2.00]0.00 [0.00, 0.00]1.00 [1.00, 1.00]
real / targetP01control0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / targetP01candidate0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / broader_familyP01control0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 20.00 [0.00, 0.00]1.33 [1.00, 2.00]0.67 [0.00, 1.00]
real / broader_familyP01candidate0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 20.67 [0.00, 1.00]1.33 [1.00, 2.00]0.00 [0.00, 0.00]
real / targetP02control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.00 [0.00, 0.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
real / targetP02candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / broader_familyP02control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.00 [0.00, 0.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
real / broader_familyP02candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / targetP03control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.00 [0.00, 0.00]0.00 [0.00, 0.00]1.00 [1.00, 1.00]
real / targetP03candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.67 [0.00, 1.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]
real / broader_familyP03control0.33 [0.00, 1.00] / 20.33 [0.00, 1.00] / 20.00 [0.00, 0.00]0.00 [0.00, 0.00]1.33 [1.00, 2.00]
real / broader_familyP03candidate0.67 [0.00, 1.00] / 20.67 [0.00, 1.00] / 20.67 [0.00, 1.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]
real / targetP04control0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / targetP04candidate0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / broader_familyP04control0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / broader_familyP04candidate0.00 [0.00, 0.00] / 00.00 [0.00, 0.00] / 00.00 [0.00, 0.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / targetP05control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.33 [0.00, 1.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
real / targetP05candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 11.00 [1.00, 1.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / broader_familyP05control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.33 [0.00, 1.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
real / broader_familyP05candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 11.00 [1.00, 1.00]0.00 [0.00, 0.00]0.00 [0.00, 0.00]
real / targetX01control0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.33 [0.00, 1.00]0.00 [0.00, 0.00]0.67 [0.00, 1.00]
real / targetX01candidate0.00 [0.00, 0.00] / 10.00 [0.00, 0.00] / 10.67 [0.00, 1.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]
real / broader_familyX01control0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 21.00 [1.00, 1.00]0.00 [0.00, 0.00]1.00 [1.00, 1.00]
real / broader_familyX01candidate0.00 [0.00, 0.00] / 20.00 [0.00, 0.00] / 21.67 [1.00, 2.00]0.00 [0.00, 0.00]0.33 [0.00, 1.00]

Candidate target defects remaining

The unchanged D3 structural classifier was applied after measurement. Classes overlap; distinct cases are counted once per class across repeats, and pooled attempts count each occurrence. Refusals and comparison-unverified answers have no inferred semantic defect class and remain separately accounted. A generic row mismatch does not isolate its cause.

DatasetClassDistinct casesPooled attemptsFamilies
freshoutput_projection_mismatch22P02, P05
freshplanner_error_unresolved11P03
freshrow_set_size_mismatch47P01, P02, P04, P05
freshsql_validation_boundary45P03, P04

fresh target outcome accounting: `{"comparison_unverified": {"distinct_cases": 10, "pooled_attempts": 22}, "correct_completed": {"distinct_cases": 2, "pooled_attempts": 5}, "refusal_unverified": {"distinct_cases": 4, "pooled_attempts": 8}, "silent_wrong": {"distinct_cases": 4, "pooled_attempts": 7}, "task_error": {"distinct_cases": 5, "pooled_attempts": 6}}`. Unclassified wrong/error attempts: 0.

DatasetClassDistinct casesPooled attemptsFamilies
realoutput_projection_mismatch25P05, X01
realrow_set_size_mismatch37P03, P05, X01
realsql_validation_boundary22P03, X01

real target outcome accounting: `{"comparison_unverified": {"distinct_cases": 1, "pooled_attempts": 3}, "silent_wrong": {"distinct_cases": 3, "pooled_attempts": 7}, "task_error": {"distinct_cases": 2, "pooled_attempts": 2}}`. Unclassified wrong/error attempts: 0.

Outcome transitions against control

Every matched-case Cartesian control-repeat × candidate-repeat pair has equal weight (nine pairings). Counts sum to one repeat’s denominator for the named set. These are outcome movements, not proven corrections.

Dataset / setControl outcomeCandidate outcomeMean count
fresh / wholecomparison_unverifiedcomparison_unverified15.22
fresh / wholecomparison_unverifiedcorrect_completed0.22
fresh / wholecomparison_unverifiedrefusal_unverified3.89
fresh / wholecomparison_unverifiedsilent_wrong1.33
fresh / wholecomparison_unverifiedtask_error0.67
fresh / wholecorrect_completedcomparison_unverified0.56
fresh / wholecorrect_completedcorrect_completed0.33
fresh / wholecorrect_completedrefusal_unverified1.00
fresh / wholecorrect_completedsilent_wrong0.33
fresh / wholecorrect_completedtask_error0.11
fresh / wholerefusal_unverifiedcomparison_unverified3.44
fresh / wholerefusal_unverifiedcorrect_completed0.44
fresh / wholerefusal_unverifiedrefusal_unverified27.67
fresh / wholerefusal_unverifiedsilent_wrong3.44
fresh / wholerefusal_unverifiedtask_error4.00
fresh / wholesilent_wrongcomparison_unverified2.00
fresh / wholesilent_wrongrefusal_unverified7.11
fresh / wholesilent_wrongsilent_wrong18.11
fresh / wholesilent_wrongtask_error2.78
fresh / wholetask_errorcomparison_unverified7.11
fresh / wholetask_errorcorrect_completed1.67
fresh / wholetask_errorrefusal_unverified5.67
fresh / wholetask_errorsilent_wrong3.78
fresh / wholetask_errortask_error11.11
fresh / targetcomparison_unverifiedcomparison_unverified0.89
fresh / targetcomparison_unverifiedrefusal_unverified0.11
fresh / targetcomparison_unverifiedsilent_wrong0.11
fresh / targetcomparison_unverifiedtask_error0.22
fresh / targetcorrect_completedcorrect_completed0.22
fresh / targetcorrect_completedtask_error0.11
fresh / targetrefusal_unverifiedrefusal_unverified0.67
fresh / targetsilent_wrongrefusal_unverified0.11
fresh / targetsilent_wrongsilent_wrong0.22
fresh / targettask_errorcomparison_unverified6.44
fresh / targettask_errorcorrect_completed1.44
fresh / targettask_errorrefusal_unverified1.78
fresh / targettask_errorsilent_wrong2.00
fresh / targettask_errortask_error1.67
fresh / restcomparison_unverifiedcomparison_unverified14.33
fresh / restcomparison_unverifiedcorrect_completed0.22
fresh / restcomparison_unverifiedrefusal_unverified3.78
fresh / restcomparison_unverifiedsilent_wrong1.22
fresh / restcomparison_unverifiedtask_error0.44
fresh / restcorrect_completedcomparison_unverified0.56
fresh / restcorrect_completedcorrect_completed0.11
fresh / restcorrect_completedrefusal_unverified1.00
fresh / restcorrect_completedsilent_wrong0.33
fresh / restrefusal_unverifiedcomparison_unverified3.44
fresh / restrefusal_unverifiedcorrect_completed0.44
fresh / restrefusal_unverifiedrefusal_unverified27.00
fresh / restrefusal_unverifiedsilent_wrong3.44
fresh / restrefusal_unverifiedtask_error4.00
fresh / restsilent_wrongcomparison_unverified2.00
fresh / restsilent_wrongrefusal_unverified7.00
fresh / restsilent_wrongsilent_wrong17.89
fresh / restsilent_wrongtask_error2.78
fresh / resttask_errorcomparison_unverified0.67
fresh / resttask_errorcorrect_completed0.22
fresh / resttask_errorrefusal_unverified3.89
fresh / resttask_errorsilent_wrong1.78
fresh / resttask_errortask_error9.44
real / wholecomparison_unverifiedcomparison_unverified19.78
real / wholecomparison_unverifiedcorrect_completed1.33
real / wholecomparison_unverifiedrefusal_unverified2.78
real / wholecomparison_unverifiedsilent_wrong1.44
real / wholecomparison_unverifiedtask_error0.67
real / wholecorrect_completedcomparison_unverified0.33
real / wholecorrect_completedcorrect_completed1.00
real / wholecorrect_completedrefusal_unverified0.22
real / wholecorrect_completedsilent_wrong0.11
real / wholerefusal_unverifiedcomparison_unverified3.89
real / wholerefusal_unverifiedcorrect_completed0.33
real / wholerefusal_unverifiedrefusal_unverified21.22
real / wholerefusal_unverifiedsilent_wrong4.22
real / wholerefusal_unverifiedtask_error0.33
real / wholesilent_wrongcomparison_unverified2.89
real / wholesilent_wrongcorrect_completed0.11
real / wholesilent_wrongrefusal_unverified2.89
real / wholesilent_wrongsilent_wrong20.00
real / wholesilent_wrongtask_error1.11
real / wholetask_errorcomparison_unverified1.44
real / wholetask_errorcorrect_completed0.22
real / wholetask_errorrefusal_unverified0.56
real / wholetask_errorsilent_wrong3.56
real / wholetask_errortask_error10.56
real / targetcomparison_unverifiedcomparison_unverified0.33
real / targetsilent_wrongsilent_wrong0.56
real / targetsilent_wrongtask_error0.11
real / targettask_errorcomparison_unverified0.67
real / targettask_errorsilent_wrong1.78
real / targettask_errortask_error0.56
real / restcomparison_unverifiedcomparison_unverified19.44
real / restcomparison_unverifiedcorrect_completed1.33
real / restcomparison_unverifiedrefusal_unverified2.78
real / restcomparison_unverifiedsilent_wrong1.44
real / restcomparison_unverifiedtask_error0.67
real / restcorrect_completedcomparison_unverified0.33
real / restcorrect_completedcorrect_completed1.00
real / restcorrect_completedrefusal_unverified0.22
real / restcorrect_completedsilent_wrong0.11
real / restrefusal_unverifiedcomparison_unverified3.89
real / restrefusal_unverifiedcorrect_completed0.33
real / restrefusal_unverifiedrefusal_unverified21.22
real / restrefusal_unverifiedsilent_wrong4.22
real / restrefusal_unverifiedtask_error0.33
real / restsilent_wrongcomparison_unverified2.89
real / restsilent_wrongcorrect_completed0.11
real / restsilent_wrongrefusal_unverified2.89
real / restsilent_wrongsilent_wrong19.44
real / restsilent_wrongtask_error1.00
real / resttask_errorcomparison_unverified0.78
real / resttask_errorcorrect_completed0.22
real / resttask_errorrefusal_unverified0.56
real / resttask_errorsilent_wrong1.78
real / resttask_errortask_error10.00

Provider identity and provenance

Requested zai model: glm-5.2; temperature 0, seed 42, thinking disabled, timeout 25 seconds. Defaults remain in generation; verifier and repair are off. This table covers generator dispatches only. Returned response.model values are retained for completed generator calls and linked to attempt job IDs. They are provider metadata, not independent serving attestation. Missing links stay explicit. Failed calls and calls without terminal records do not attest a served model.

RunMeasured revisionGenerator completed / failed / no terminal recordGenerator returned model IDsAttempts without linked generator calls
control-fresh-r1`2513a7186df76693dd4085a38a4aca865fcefa1a`248 / 0 / 0{"glm-5.3": 248}1
control-fresh-r2`2513a7186df76693dd4085a38a4aca865fcefa1a`257 / 0 / 0{"glm-5.3": 257}0
control-fresh-r3`2513a7186df76693dd4085a38a4aca865fcefa1a`251 / 0 / 0{"glm-5.3": 251}0
candidate-fresh-r1`d1d63b5eb4d9ccf987992be7e7d714ff62509370`264 / 0 / 0{"glm-5.3": 264}0
candidate-fresh-r2`d1d63b5eb4d9ccf987992be7e7d714ff62509370`261 / 0 / 0{"glm-5.3": 261}0
candidate-fresh-r3`d1d63b5eb4d9ccf987992be7e7d714ff62509370`255 / 1 / 0{"glm-5.3": 255}0
control-real-r1`2513a7186df76693dd4085a38a4aca865fcefa1a`179 / 1 / 0{"glm-5.3": 179}2
control-real-r2`b24c1c3f0381bb326eef039b6428e0893887c04a`174 / 0 / 0{"glm-5.3": 174}2
control-real-r3`b24c1c3f0381bb326eef039b6428e0893887c04a`172 / 0 / 0{"glm-5.3": 172}2
candidate-real-r1`d1d63b5eb4d9ccf987992be7e7d714ff62509370`171 / 0 / 0{"glm-5.3": 171}2
candidate-real-r2`d1d63b5eb4d9ccf987992be7e7d714ff62509370`172 / 1 / 0{"glm-5.3": 172}2
candidate-real-r3`d1d63b5eb4d9ccf987992be7e7d714ff62509370`185 / 0 / 0{"glm-5.3": 185}2

A separate auxiliary wrapper retained 596 completed invocations. Its model/provider labels come from execution configuration, not returned response identity. Returned auxiliary identities and individual internal failed/fallback dispatches were not separately retained and remain unverified. These records cannot establish a complete provider-dispatch inventory. `auxiliary-provenance.json` records per-run counts and source hashes; `auxiliary/` retains the structural call/job linkage with explicitly named configured-model fields. No missing calls or identities were reconstructed.

New-run chronology: queue-to-entry includes process launch and scheduling, so it is an upper bound on lock wait. Start below is driver entry, before host admission; finish includes owned-process cleanup.

RunDriver entry (UTC)Finished (UTC)Queue-to-entry (s)Listener port
candidate-fresh-r12026-09-12T04:33:13.853122+00:002026-09-12T04:50:03.737539+00:00815.378700
candidate-fresh-r22026-09-12T05:03:24.570141+00:002026-09-12T05:20:08.828103+00:00800.488700
candidate-fresh-r32026-09-12T05:32:47.614021+00:002026-09-12T05:52:02.112767+00:00758.348700
control-real-r22026-09-12T03:31:00.347154+00:002026-09-12T03:44:42.678897+00:000.248700
control-real-r32026-09-12T04:06:39.189695+00:002026-09-12T04:19:38.077654+00:001316.118700
candidate-real-r12026-09-12T16:19:53.941708+00:002026-09-12T16:32:18.839200+00:000.228700
candidate-real-r22026-09-12T16:32:19.468932+00:002026-09-12T16:45:11.399448+00:000.238700
candidate-real-r32026-09-12T16:45:12.011562+00:002026-09-12T16:58:30.533776+00:000.218700

Historical D3 fresh control repeats and real control repeat 1 used revision 2513a718. Two new real control repeats use the D3 tip; its documented changes are obligation-recording and sampling-label telemetry. The D3 offline replay reproduced outcomes but did not remeasure telemetry cost. The three candidate repeats per dataset use one frozen revision. Corrected corpus and evaluator hashes agree within each dataset, as do requested generator/verifier configurations. DCLASS containment and telemetry wrappers differ from historical D3; their separate source hashes are retained. The DCLASS harness sets `LORE_GRANOLA_INDEX_AUTO_INIT=0` to prevent shared-index initialization writes. This containment setting differs from historical wrapper behavior; it is explicitly retained as a method difference. Fresh-v1 was already exposed in D2/D3, and historical failures plus question shapes informed this root fix. There is no claim of an unseen holdout or a sole-cause production effect.

New measurements hold `/tmp/lore-eval.lock` across admission, run and cleanup, use two workers and application ports 8700–8799, and retain queue-to-entry wait evidence. This does not eliminate unrelated shared-host activity. Frozen application time is 2026-09-06 12:00 UTC, with Europe/London business periods. Per-question p95 includes every attempt with a retained feature-completion time; failed-call timings remain in traces.

The session resumed after five new runs had completed, retaining all three candidate fresh repeats and both supplemental real controls. The final fresh receipt records successful completion and cleanup although its controller sequence event was not written. Receipt validation allowed the three remaining real repeats to proceed without rerunning any completed measurement. The interruption adds a substantial time gap before candidate real measurements; provider and host conditions across that gap are not controlled. `resume-receipt.json` preserves the recovery evidence.

Checks, limitations and review

Check evidence: `{"curated": {"errors": 0, "failed": 0, "passed": 636, "skipped": 88}, "errors": 0, "exit_code": 0, "failed": 0, "files": 46, "harness": {"errors": 0, "failed": 0, "passed": 43, "receipt": "checks/harness-final-checks.json", "result_line": "43 passed in 13.17s", "ruff": "All checks passed!"}, "missing_inherited": ["tests/js/bi-terminal-rendering.test.mjs", "tests/test_answer_contract_role_promotion.py", "tests/test_federated_semantic_refusal.py", "tests/test_insights_cold_schema_context.py", "tests/test_insights_federated_followup_contract.py", "tests/test_insights_n01_definition_schema.py", "tests/test_insights_v16_aggregate_group_scope.py", "tests/test_insights_v16_rank_words.py", "tests/test_insights_v17_distinct_snapshot_sum.py", "tests/test_insights_v17_explicit_output.py", "tests/test_insights_v17_membership_predicates.py", "tests/test_insights_v17_output_recovery_native.py", "tests/test_insights_v18_membership_output.py", "tests/test_insights_v18_money_guard_recovery.py", "tests/test_insights_v8_chain_recovery.py", "tests/test_insights_v8_context_contract.py", "tests/test_insights_v8_ct_input_regressions.py", "tests/test_insights_v8_intent_shape_diagnostics.py", "tests/test_insights_v9_contract_review.py"], "missing_inherited_files": 19, "passed": 1045, "receipt_tests": {"errors": 0, "failed": 0, "passed": 8, "result_line": "8 passed in 13.14s", "ruff": "All checks passed!"}, "result_lines": ["1045 passed, 92 skipped, 23 warnings in 191.92s (0:03:11)"], "skipped": 92, "wall_seconds": 202.26604114286602}`.

Zero observed curated violations covers present executed tests only. Skipped and absent inherited tests are unverified. The ratified staff/house and Pipedrive-only identity entries are applied unchanged where relevant; owner dependencies remain pending. In particular X01 people comparisons and P02 listings can depend on these meanings. The benchmark uses real question shapes against a fixture and corrected oracle; it does not establish correctness on live customer data. Family-bootstrap intervals contain only four target real clusters and omit run-level uncertainty. The real-target correctness interval [0,0] is degenerate conditional on the observed all-zero counts; it does not establish population equivalence or serving stability. Each four-question real-target nearest-rank p95 equals that repeat’s maximum, rather than a well-estimated population tail.

Independent review and correction ledger

Three separate foreground Codex CLI processes reviewed checkpoint `5e7e25b6` with requested model `gpt-6-astra`, reasoning `high`, full-access sandbox and strict no-modification instructions. Each exited successfully and detected zero changed inputs among 6,507 checked files. The additional inventory covering 568 report, PR, source, corpus and full/compact measurement inputs also detected zero changes across all three reviews. Prompts, replies and receipts are retained; the CLI banner is configuration evidence, not independent serving attestation.

ReviewVerdictFinding → action
Arithmetic and intervalsPASS; experiment remains FAILIndependently reproduced 1,338 raw attempts, 60 bootstrap intervals, six transition tables, family metrics and frozen gates. No numerical correction.
Leakage, model identity and noiseFailed experiment conclusion supported; reporting correction requiredGenerator provenance did not describe 596 auxiliary invocations → scoped the model table to generator dispatches, exported auxiliary structural records with configured-model labels, and disclosed missing auxiliary returned identities/internal dispatch detail. No confirmed new disclosure or post-freeze application tuning was found.
Unsupported wordingNo consequential numerical contradiction; presentation corrections requiredPrecision means used only defined repeats → added n and explicit undefined handling. The real remaining-defect table lacked a header → repeated the header. Added explicit limits for the degenerate [0,0] interval and four-question p95, which equals the sample maximum.

The reviewed results SHA-256 is `1d79d83fb1a8fad395729d303a95003944e361650c01526fbf8ccc4beba58a11`. All quality metrics, intervals, transitions, thresholds, membership, raw records and application code remain unchanged after review. Corrections affect reporting and structural provenance only; no missing auxiliary calls or model identities were reconstructed and no measurement was rerun. These reporting corrections were applied after the three reviews and checked locally; no fourth independent review is claimed.

Before those reviews, the resumed session also corrected the refusal-partition description, an environment-setting name and an inherited generator model-source label; added run chronology and the interrupted-session disclosure; and foregrounded the target wrong-answer increase. The complete compact JSONL records are now committed despite the repository-wide JSONL ignore rule. These were reporting/retention corrections, not application tuning.

Independent reviews validated structural outcome fields, timing and retained checks. They did not perform new semantic answer scoring, execute the application tests, verify live serving identity or validate sibling integration. The absence of a classified schema-qualification error does not establish complete semantic diagnosis or correctness of unverified answers.

Expected merge order: D1+D2+D3 (PR #104), then DCLASS, then sibling F01. The localized schema-loader correction establishes the shared boundary before broader family-context work; F01 must preserve it when resolving overlap. Combined sibling behavior is unverified. No application merge or deployment is part of this report.

Full retained records: `/tmp/lore-goal3-eval/goal-dclass/`; historical controls: `/tmp/lore-goal3-eval/goal-d3/`. Compact records: `docs/evidence/insights-goal-dclass-2026-09-12/`. Recompute with `PYTHONDONTWRITEBYTECODE=1 /home/stevan/dev/lore/.venv/bin/python scripts/insights_dclass_report.py --compact`. Compact evidence excludes question text, SQL, rows and free-text reasoning. HTML publication is public: random naming and noindex discourage discovery and indexing but do not provide access control.