###Background
Theexperimentfrontmatterschemaisinbettershapethanbefore:analysis_type,tags,andnotifyarealreadyparsedinpkg/workflow/compiler_experiments.goandsurfacedinactions/setup/js/pick_experiment.cjs.However,theend-to-endanalyticspipelineisstillincompletebecausethecurrentrunartifactpath/tmp/gh-aw/agent/experiments/state.jsonisnotpopulatedconsistentlyinthisenvironment,andexperimentreporting/auditsurfacesremainfragmented.
###VerifiedArea1Findings
ViewDetails
Thefield-presencecheckoutcomefortherequestedfieldsiseffectively:
-analysis_type:presentincompilerparsingandinpick_experiment.cjsreportingoutput
-tags:presentincompilerparsingandinpick_experiment.cjsreportingoutput
-notify:presentincompilerparsingandinpick_experiment.cjsreportingoutput
Evidencehighlights:
-pkg/workflow/compiler_experiments.goparsesanalysis_type,tags,andnotify
-actions/setup/js/pick_experiment.cjsdocumentsandrendersanalysis_type,tags,andnotifyinstep-summarydetailoutput
Becausethesefieldsareimplemented,thisissueshouldnotproposeadditionalschemaworkforthem.Instead,itshouldfocusonsurfacing,aggregation,analyticsautomation,andauditability.
###ProblemStatement
Experimentscanbeconfigured,assigned,andpartiallysurfaced,butthereisnotyetacompleteclosed-loopsystemthat:
1.capturesassignment+outcome+cost/latencysignalsinastableartifactschema,
2.aggregatesresultsacrossrunsautomatically,
3.computessignificanceandwinnerrecommendations,and
4.exposesexperimentstateconsistentlyinghawauditandOTELtraces.
###ProposedImprovements
####1.Stableexperimentrunartifactcontract
Defineanddocumentacanonicalexperiments/state.jsonpayloadwrittenoneveryexperimentalrun,forexample:
{
"workflow":"outcome-collector",
"run_id":31253134640,
"timestamp":"2026-08-08T00:00:00Z",
"experiments":[
{
"name":"max_turns_campaign",
"variant":"balanced",
"assignment_method":"round_robin_cache",
"analysis_type":"mann_whitney",
"tags":["latency","cost"],
"metrics":{
"run_duration_seconds":412,
"ai_credits_estimate":1.72,
"output_length_chars":6842,
"success":true,
"noop":false,
"guardrails":{
"malformed_report":false,
"empty_output":false
}
}
}
]
}
Concretework:
-alwaysemitthefilewhenexperiments:isactive
-validateitagainstaJSONschemainCI
-includebothassignmentmetadataandobservedmetricpayloads
-captureexplicitguardrailevaluationresults,notjustrawvalues
####2.Daily/weeklyreportingworkflow
Addadedicatedreportworkflowthat:
-downloadsrecentworkflowrunartifactsforexperiments
-aggregatesperexperiment/variantsamplesize,mean,variance,andguardrailcounts
-computessignificancewhenp<0.05
-rendersamarkdownsummaryandanASCIIcomparisonchartartifact
-poststhecurrentwinnerandconfidenceleveltoadiscussionorlinkedissue
Suggestedreportoutputshape:
###DailyExperimentReport
|Experiment|Variant|n|Meanefficiency|Failurerate|Currentleader|
|---|---:|---:|---:|---:|---|
|max_turns_campaign|tight|8|0.71|0.0%||
|max_turns_campaign|balanced|9|0.83|0.0%|✅balanced|
|max_turns_campaign|extended|8|0.78|0.0%||
<details><summary>ViewDetails</summary>
-Test:Mann-WhitneyU
-p-value:0.031
-Guardrails:allpassing
-Recommendation:continueuntil21runs/variant,butbalancediscurrentlyleading
</details>
Implementationnotes:
-useanalysis_typefromfrontmattertochoosethestatisticaltest
-computeper-variantrunningmeansandpooledcomparisonsincrementally
-persistrollupsincache-memorytoavoidreprocessingallhistoricalrunseachtime
####3.AuditandOTELintegration
Extendobservabilitysoeveryexperimentassignmentisqueryable:
-addOTELspanattributes:experiment.name,experiment.variant,experiment.assignment_method,experiment.issue
-includeexperimentmetadatainghawauditoutputforeachrunandstep
-supportfilteringauditresultsbyvarianttocomparefailuremodes
-addastepsummaryblockfrompick_experiment.cjsthatincludesassignment,samplecounts,andlinkedtrackingissue/discussion
ExampleauditUX:
ghawauditoutcome-collector.lock.yml--experimentmax_turns_campaign
Run31253134640variant=balancedsuccess=trueduration=412smalformed_report=false
Run31251000000variant=tightsuccess=trueduration=366smalformed_report=false
Run31249000000variant=extendedsuccess=falseduration=900smalformed_report=true
####4.Notificationcompletionloop
Thenotifyfieldexists,butsignificance-awarenotificationflowshouldbeautomated:
-whenmin_samplesisreachedandsignificancecriteriapass,emitamachine-generatedrecommendation
-posttoconfigurednotify.issueornotify.discussion
-includewinner,effectsize,guardrailstatus,andstop/continuerecommendation
-suppressnoisyre-postingunlesstheleaderchangesorconfidencemateriallyincreases
###AcceptanceCriteria
-[]Everyrunwithactiveexperimentswritesavalidatedexperiments/state.jsonartifact
-[]Areportworkflowproducesdaily/weeklyaggregatesummariesacrossvariants
-[]Statisticaltestselectionisdrivenbyanalysis_type
-[]ghawauditcandisplayandfilterexperimentassignmentsandoutcomes
-[]OTELspansincludeexperimentnameandvariantattributes
-[]Configurednotifytargetsreceivesignificance/winnerupdatesautomatically
###References
-pkg/workflow/compiler_experiments.go
-actions/setup/js/pick_experiment.cjs
-/tmp/gh-aw/agent/experiments/state.json(expectedrunartifactpath)
-Parentcampaignissue:#aw_abinfra1
Generated by 🧪 Daily A/B Testing Advisor · gpt54 · 17.4 AIC · ⌖ 5.96 AIC · ⊞ 8.3K · ◷
###Background
Theexperimentfrontmatterschemaisinbettershapethanbefore:
analysis_type,tags,andnotifyarealreadyparsedinpkg/workflow/compiler_experiments.goandsurfacedinactions/setup/js/pick_experiment.cjs.However,theend-to-endanalyticspipelineisstillincompletebecausethecurrentrunartifactpath/tmp/gh-aw/agent/experiments/state.jsonisnotpopulatedconsistentlyinthisenvironment,andexperimentreporting/auditsurfacesremainfragmented.###VerifiedArea1Findings
ViewDetails
Thefield-presencecheckoutcomefortherequestedfieldsiseffectively:
-
analysis_type:presentincompilerparsingandinpick_experiment.cjsreportingoutput-
tags:presentincompilerparsingandinpick_experiment.cjsreportingoutput-
notify:presentincompilerparsingandinpick_experiment.cjsreportingoutputEvidencehighlights:
-
pkg/workflow/compiler_experiments.goparsesanalysis_type,tags,andnotify-
actions/setup/js/pick_experiment.cjsdocumentsandrendersanalysis_type,tags,andnotifyinstep-summarydetailoutputBecausethesefieldsareimplemented,thisissueshouldnotproposeadditionalschemaworkforthem.Instead,itshouldfocusonsurfacing,aggregation,analyticsautomation,andauditability.
###ProblemStatement
Experimentscanbeconfigured,assigned,andpartiallysurfaced,butthereisnotyetacompleteclosed-loopsystemthat:
1.capturesassignment+outcome+cost/latencysignalsinastableartifactschema,
2.aggregatesresultsacrossrunsautomatically,
3.computessignificanceandwinnerrecommendations,and
4.exposesexperimentstateconsistentlyin
ghawauditandOTELtraces.###ProposedImprovements
####1.Stableexperimentrunartifactcontract
Defineanddocumentacanonical
experiments/state.jsonpayloadwrittenoneveryexperimentalrun,forexample:{ "workflow":"outcome-collector", "run_id":31253134640, "timestamp":"2026-08-08T00:00:00Z", "experiments":[ { "name":"max_turns_campaign", "variant":"balanced", "assignment_method":"round_robin_cache", "analysis_type":"mann_whitney", "tags":["latency","cost"], "metrics":{ "run_duration_seconds":412, "ai_credits_estimate":1.72, "output_length_chars":6842, "success":true, "noop":false, "guardrails":{ "malformed_report":false, "empty_output":false } } } ] }Concretework:
-alwaysemitthefilewhen
experiments:isactive-validateitagainstaJSONschemainCI
-includebothassignmentmetadataandobservedmetricpayloads
-captureexplicitguardrailevaluationresults,notjustrawvalues
####2.Daily/weeklyreportingworkflow
Addadedicatedreportworkflowthat:
-downloadsrecentworkflowrunartifactsforexperiments
-aggregatesperexperiment/variantsamplesize,mean,variance,andguardrailcounts
-computessignificancewhen
p<0.05-rendersamarkdownsummaryandanASCIIcomparisonchartartifact
-poststhecurrentwinnerandconfidenceleveltoadiscussionorlinkedissue
Suggestedreportoutputshape:
Implementationnotes:
-use
analysis_typefromfrontmattertochoosethestatisticaltest-computeper-variantrunningmeansandpooledcomparisonsincrementally
-persistrollupsincache-memorytoavoidreprocessingallhistoricalrunseachtime
####3.AuditandOTELintegration
Extendobservabilitysoeveryexperimentassignmentisqueryable:
-addOTELspanattributes:
experiment.name,experiment.variant,experiment.assignment_method,experiment.issue-includeexperimentmetadatain
ghawauditoutputforeachrunandstep-supportfilteringauditresultsbyvarianttocomparefailuremodes
-addastepsummaryblockfrom
pick_experiment.cjsthatincludesassignment,samplecounts,andlinkedtrackingissue/discussionExampleauditUX:
####4.Notificationcompletionloop
The
notifyfieldexists,butsignificance-awarenotificationflowshouldbeautomated:-when
min_samplesisreachedandsignificancecriteriapass,emitamachine-generatedrecommendation-posttoconfigured
notify.issueornotify.discussion-includewinner,effectsize,guardrailstatus,andstop/continuerecommendation
-suppressnoisyre-postingunlesstheleaderchangesorconfidencemateriallyincreases
###AcceptanceCriteria
-[]Everyrunwithactiveexperimentswritesavalidated
experiments/state.jsonartifact-[]Areportworkflowproducesdaily/weeklyaggregatesummariesacrossvariants
-[]Statisticaltestselectionisdrivenby
analysis_type-[]
ghawauditcandisplayandfilterexperimentassignmentsandoutcomes-[]OTELspansincludeexperimentnameandvariantattributes
-[]Configured
notifytargetsreceivesignificance/winnerupdatesautomatically###References
-
pkg/workflow/compiler_experiments.go-
actions/setup/js/pick_experiment.cjs-
/tmp/gh-aw/agent/experiments/state.json(expectedrunartifactpath)-Parentcampaignissue:#aw_abinfra1