###🧪ExperimentCampaign:outcome-collector
Workflowfile:.github/workflows/outcome-collector.md
Selecteddimension:max_turns
Triggeredby:ab-testing-advisoron2026-08-08
###Background
outcome-collectorisaperiodicreportingworkflowthatturnsprecomputedsafe-outputoutcomedataintoanexecutive-readyissuewithactionitems,lifecyclehealthlabels,andadetailedmetricsappendix.Ichosemax_turnsbecausethisworkflowhasalong,highlystructuredpromptbutaboundedtasksurface;itisagoodcandidatefortestingwhethertighterturnbudgetscanpreservereportqualitywhilereducingcostandlatency.
###Hypothesis
**Nullhypothesis(H0):**Reducingorincreasingtheturnbudgetdoesnotimprovereportcompletionqualityorefficiencycomparedtothebaselineturnbudget.
**Alternativehypothesis(H1):**Amoderateturnbudgetimprovessuccessfulfirst-passreportgenerationefficiencybyatleast15%onAI-credit-adjustedruntimewithoutincreasingformattingdefects,emptyreports,ornoopmisuse.
ViewDetails
###ExperimentConfiguration
Addthefollowingexperiments:blocktotheworkflowfrontmatter(usetherichobjectformsoallmetadataisself-documenting):
experiments:
max_turns_campaign:
variants:[tight,balanced,extended]
description:"Measurewhetheroutcomereportqualitycanbepreservedwithasmallerorlargeragentturnbudget."
hypothesis:"H0:nochangeinreportcompletionefficiencyorquality.H1:thebalancedvariantreducesAI-credit-adjustedruntimeby>=15%versusbaselinewithoutincreasingreportdefects."
metric:report_completion_efficiency
secondary_metrics:[run_duration_minutes,output_length_chars,noop_rate]
guardrail_metrics:
-name:malformed_report_rate
direction:min
threshold:0.05
-name:empty_output_rate
direction:min
threshold:0.01
-name:run_failure_rate
direction:min
threshold:0.05
min_samples:21
weight:[34,33,33]
start_date:"2026-08-08"
issue:<this_issue_number>
Variantdescriptions:
-tight:Captheagentatasmallnumberofturnsandbiastheprompttowarddirectsynthesisfromtheprecomputedfiles.
-balanced:Useamoderateturnbudgetintendedtopreservecurrentqualitywithlessexploratoryback-and-forth.
-extended:Allowalargerturnbudgetformoreiterativecheckingandricherper-workflowreasoning.
###WorkflowChangesRequired
Listtheexactchangesneededintheworkflowmarkdownbodytoimplementtheexperimentusinghandlebarsconditionalblocks.Alwayscompareagainstaspecificvariantvalue—thecorrectsyntaxis{{#ifexperiments.<name>=="<variant>"}}...{{else}}...{{/if}}.
Concretefrontmatter/bodydiff:
---a/.github/workflows/outcome-collector.md
+++b/.github/workflows/outcome-collector.md
@@
timeout-minutes:20
+experiments:
+max_turns_campaign:
+variants:[tight,balanced,extended]
+description:"Measurewhetheroutcomereportqualitycanbepreservedwithasmallerorlargeragentturnbudget."
+hypothesis:"H0:nochangeinreportcompletionefficiencyorquality.H1:thebalancedvariantreducesAI-credit-adjustedruntimeby>=15%versusbaselinewithoutincreasingreportdefects."
+metric:report_completion_efficiency
+secondary_metrics:[run_duration_minutes,output_length_chars,noop_rate]
+guardrail_metrics:
+-name:malformed_report_rate
+direction:min
+threshold:0.05
+-name:empty_output_rate
+direction:min
+threshold:0.01
+-name:run_failure_rate
+direction:min
+threshold:0.05
+min_samples:21
+weight:[34,33,33]
+start_date:"2026-08-08"
+issue:<this_issue_number>
@@
-4.Otherwise,createareportissuewiththesummary
+4.Otherwise,createareportissuewiththesummary
+
+{{#ifexperiments.max_turns_campaign=="tight"}}
+###Turn-budgetvariantinstructions
+
+-Workinasinglesynthesispasswhenpossible.
+-Donotperformoptionalre-checkloopsafterthefirstvalidreportdraft.
+-Preferconciseactionitemsoverexpandednarrative.
+{{else}}
+{{#ifexperiments.max_turns_campaign=="extended"}}
+###Turn-budgetvariantinstructions
+
+-Performoneextraself-checkpassonexecutive-summary/referencealignmentbeforepublishing.
+-Expandper-workflowreasoningwhenclassifyingstuckvs.underdefinedworkflows.
+-Ifdataqualitysignalsareambiguous,spendoneextraturnreconcilingthembeforedeciding.
+{{else}}
+###Turn-budgetvariantinstructions
+
+-Useonevalidationpassafterdraftingtheexecutivesection.
+-Keepper-workflowreasoningconcisebutexplicit.
+{{/if}}
+{{/if}}
###SuccessMetrics
| Metric |
Type |
Target |
report_completion_efficiency |
Primary |
≥15%improvementforwinningvariantvs.baseline-equivalentcontrol |
run_duration_minutes |
Secondary |
Downwardtrendwithoutqualityloss |
output_length_chars |
Secondary |
Stablewithin±20%unlessqualityclearlyimproves |
noop_rate |
Secondary |
Noincreasebeyondbaselinenoise |
malformed_report_rate |
Guardrail |
Mustremain≤5% |
empty_output_rate |
Guardrail |
Mustremain≤1% |
run_failure_rate |
Guardrail |
Mustremain≤5% |
###StatisticalDesign
-Variants:tight,balanced,extended
-Assignment:Round-robinviagh-awexperimentsruntime(cache-based)
-Minimumrunspervariant:21
-Expectedexperimentduration:~63daysatthecurrentevery-3-daysschedulefor21runspervariant;ifmanualdispatchisusedperiodically,durationshortensproportionally
-Analysisapproach:Mann-WhitneyUtestonefficiencyscore,withguardrailproportiontestsformalformed/empty/failurerates
###ImplementationSteps
-[]Addexperiments:sectiontofrontmatter
-[]Addconditionalblockstoworkflowpromptbodyusing{{#ifexperiments.max_turns_campaign=="<variant>"}}(value-comparisonform—neverusetheinternal__GH_AW_EXPERIMENTS__env-varsyntax)
-[]Runghawcompileoutcome-collectortoregeneratelockfile
-[]Monitorexperimentartifactuploadedperrunto/tmp/gh-aw/agent/experiments/state.json
-[]Aftersufficientruns,analyzevariantdistributionviaworkflowrunartifacts
-[]Documentfindingsandpromotewinningvariant
###References
-A/BTestingingh-aw
-Workflowfile:.github/workflows/outcome-collector.md
###Notes
Recentvisiblerunsforoutcome-collector.lock.ymlwerepredominantlysuccessful,whichmakesitagoodcandidateforanefficiency-focusedexperimentratherthanareliabilityrescue.TheghrunlistrequestinthisenvironmentdidnotexposedurationMS,sothecampaignrecommendstrackingruntimefromavailableruntimestampsandtheexperimentartifactinstead.
###InfrastructureStatus
Area1gatecheckshowsanalysis_type,tags,andnotifyalreadyexistinbothpkg/workflow/compiler_experiments.goandactions/setup/js/pick_experiment.cjs,sothefrontmatterschemagapisclosed.Afollow-upinfrastructureissueisstillwarrantedforreporting,analytics,andauditintegrationcompleteness.
Generated by 🧪 Daily A/B Testing Advisor · gpt54 · 17.4 AIC · ⌖ 5.96 AIC · ⊞ 8.3K · ◷
###🧪ExperimentCampaign:outcome-collector
Workflowfile:
.github/workflows/outcome-collector.mdSelecteddimension:
max_turnsTriggeredby:
ab-testing-advisoron2026-08-08###Background
outcome-collectorisaperiodicreportingworkflowthatturnsprecomputedsafe-outputoutcomedataintoanexecutive-readyissuewithactionitems,lifecyclehealthlabels,andadetailedmetricsappendix.Ichosemax_turnsbecausethisworkflowhasalong,highlystructuredpromptbutaboundedtasksurface;itisagoodcandidatefortestingwhethertighterturnbudgetscanpreservereportqualitywhilereducingcostandlatency.###Hypothesis
**Nullhypothesis(H0):**Reducingorincreasingtheturnbudgetdoesnotimprovereportcompletionqualityorefficiencycomparedtothebaselineturnbudget.
**Alternativehypothesis(H1):**Amoderateturnbudgetimprovessuccessfulfirst-passreportgenerationefficiencybyatleast15%onAI-credit-adjustedruntimewithoutincreasingformattingdefects,emptyreports,ornoopmisuse.
ViewDetails
###ExperimentConfiguration
Addthefollowing
experiments:blocktotheworkflowfrontmatter(usetherichobjectformsoallmetadataisself-documenting):Variantdescriptions:
-
tight:Captheagentatasmallnumberofturnsandbiastheprompttowarddirectsynthesisfromtheprecomputedfiles.-
balanced:Useamoderateturnbudgetintendedtopreservecurrentqualitywithlessexploratoryback-and-forth.-
extended:Allowalargerturnbudgetformoreiterativecheckingandricherper-workflowreasoning.###WorkflowChangesRequired
Listtheexactchangesneededintheworkflowmarkdownbodytoimplementtheexperimentusinghandlebarsconditionalblocks.Alwayscompareagainstaspecificvariantvalue—thecorrectsyntaxis
{{#ifexperiments.<name>=="<variant>"}}...{{else}}...{{/if}}.Concretefrontmatter/bodydiff:
###SuccessMetrics
report_completion_efficiencyrun_duration_minutesoutput_length_charsnoop_ratemalformed_report_rateempty_output_raterun_failure_rate###StatisticalDesign
-Variants:
tight,balanced,extended-Assignment:Round-robinvia
gh-awexperimentsruntime(cache-based)-Minimumrunspervariant:21
-Expectedexperimentduration:~63daysatthecurrentevery-3-daysschedulefor21runspervariant;ifmanualdispatchisusedperiodically,durationshortensproportionally
-Analysisapproach:Mann-WhitneyUtestonefficiencyscore,withguardrailproportiontestsformalformed/empty/failurerates
###ImplementationSteps
-[]Add
experiments:sectiontofrontmatter-[]Addconditionalblockstoworkflowpromptbodyusing
{{#ifexperiments.max_turns_campaign=="<variant>"}}(value-comparisonform—neverusetheinternal__GH_AW_EXPERIMENTS__env-varsyntax)-[]Run
ghawcompileoutcome-collectortoregeneratelockfile-[]Monitorexperimentartifactuploadedperrunto
/tmp/gh-aw/agent/experiments/state.json-[]Aftersufficientruns,analyzevariantdistributionviaworkflowrunartifacts
-[]Documentfindingsandpromotewinningvariant
###References
-A/BTestingingh-aw
-Workflowfile:
.github/workflows/outcome-collector.md###Notes
Recentvisiblerunsfor
outcome-collector.lock.ymlwerepredominantlysuccessful,whichmakesitagoodcandidateforanefficiency-focusedexperimentratherthanareliabilityrescue.TheghrunlistrequestinthisenvironmentdidnotexposedurationMS,sothecampaignrecommendstrackingruntimefromavailableruntimestampsandtheexperimentartifactinstead.###InfrastructureStatus
Area1gatecheckshows
analysis_type,tags,andnotifyalreadyexistinbothpkg/workflow/compiler_experiments.goandactions/setup/js/pick_experiment.cjs,sothefrontmatterschemagapisclosed.Afollow-upinfrastructureissueisstillwarrantedforreporting,analytics,andauditintegrationcompleteness.