Experiment Analysis after the Design Is Frozen (2026)

By William Zhu & the InfiniSynapse Data Team · Published: 2026-08-22 · Last updated: 2026-08-24 · Last verified: 2026-08-24 · Next review: 2026-11-24 · Editorial standards · Corrections

Table of Contents

TL;DR

We evaluate these patterns at the InfiniSynapse desk on sanitized composites; sample figures on this page are illustrative, not customer uplifts.

Direct answer: Experiment analysis after freeze is a read of the assignment table first—unit, variant, time, eligibility—then outcomes, guardrails, and a memo a human signs. Lift is the last number, not the first slide.

What you'll learn: why experiment analysis starts with assignment integrity; a six-row read-order table; how ITT and exposed-only differ; how to bind metric notes; an illustrative checkout pack; and the breaks that reopen a “frozen” design.

Experiment analysis is not a dashboard refresh. It is a contract replay. If the design note is missing, label the pack exploratory and stop treating the interval as a launch order.

What Experiment Analysis Means after Freeze

Key Definition: Experiment analysis is the post-design read of a frozen assignment log against named outcomes and guardrails, producing a ship / hold / iterate memo a reviewer can audit. The freeze locks unit, primary, window, and peeking rule; it does not lock the ship decision.

Public sampling frames at the U.S. Census Bureau are a useful analogy: you do not invent a population after the field period starts. Experiment analysis inherits the same discipline. If the assignment grain changed mid-test, the freeze is already broken.

Experiment analysis sits next to the hub method in A/B test analysis. This page is narrower: what you do after the note is signed, not how you pick CUPED or write the first hypothesis.

Assignment is the object, lift is the comment

Experiment analysis that opens on conversion will shop for a winner. Open the assignment table: one unit, one variant, an assigned_at that precedes every outcome you will count. If a user appears in both arms, stop. If timestamps are missing, stop. Exploratory data analysis on balance and sample-ratio mismatch belongs here, before any lift sentence.

Claims integrity work at CMS is a reminder that a locked code is not optional decoration. Experiment analysis needs the same lock on “conversion” and “refund,” bound in a knowledge base to the outcome source—not a chat synonym that drifts by teammate.

Freeze is a protocol, not a mood

Protocol freeze in a trial sense, as the FDA treats investigational design, is the right bar even when the change is a checkout button. Experiment analysis after freeze may still find an SRM, a leaking exposure flag, or an unavailable guardrail. Those findings change the decision; they do not rewrite the primary after the fact.

If the next object is variance reduction rather than assignment order, continue in CUPED explained. If the next object is a kill metric, use guardrail metrics.

A Read-Order Framework for Frozen Tests

Every experiment analysis read should fill this table before anyone debates lift.

Read orderWhat you lockTypical sourceFailure if skipped
1. UnitUser, account, or sessionAssignment logCookie splits one person
2. VariantOne row per unitSame logDual-exposed users
3. Timeassigned_at before outcomesSame logPre-period counted as lift
4. EligibilityWho was allowed inFlag or joinIneligible “wins”
5. Primary + guardrailsOne primary, named killsOutcome + metric noteMetric shopping
6. Decision ownerHuman who shipsMemoChat as approval

Experiment analysis quality is the filled order, not the novelty of the estimator. If refunds are a guardrail and the extract has no refund grain, write “not measured.” Hiding the gap is worse than a hold.

Published extracts on data.gov carry a vintage. Experiment analysis on a dated dump needs the same stamp: file date, window, and whether later events were appended.

How Teams Compare Post-Design Reads

Teams argue tools. After freeze they should argue inclusion. Experiment analysis methods differ in which rows they keep.

Inclusion familyWorks whenBreaks when
Intent-to-treatAssignment is the policyExposure is sold as ITT
Exposed-onlyExposure is logged cleanlyExposure is inferred from clicks
Per-protocolCompliance is rare and namedNon-compliance is dropped quietly
Slice-firstPre-registered segmentsSegments invented after lift

Official statistics practice at the UK Office for National Statistics publishes the population and the exclusions. Experiment analysis should do the same in one paragraph: who was assigned, who was dropped, and why.

Intent-to-treat versus exposed-only

Classic experiment analysis is intent-to-treat: you analyze the unit you randomized. Exposed-only is a different question—did people who saw the UI convert?—and it is easy to bias. If you use it, say so, and keep the ITT line in the memo. A data agent can compute both; it cannot choose the estimand for you.

Slices that reopen the design

A slice that was not in the note is a new test. Experiment analysis may report it as exploratory. It may not promote it to the primary because the original interval looked small. That promotion is how frozen designs thaw overnight.

Tool Landscape for Assignment-First Reads

You do not need a new warehouse to finish experiment analysis. You need the assignment log, the outcome grain, and a metric note bound to those sources. A dated CSV is valid if it includes assignment time.

AI for data analysis can draft the memo. It is not NLP2SQL theater and it is not ChatBI. The human still names the primary and signs the last line.

Assignment logs you can reopen

Minimum columns: unit_id, variant, assigned_at, eligibility. Experiment analysis that lacks those four is a story. If randomization lived in another system, keep that system’s IDs in the extract. Data governance at experiment grain means one definition of “assigned,” not three Slack threads.

Metric notes bound to sources

“Conversion” is a sentence. Bind it. Experiment analysis that lets each teammate redefine conversion will produce two memos. InfiniSynapse binds a knowledge base to the data source you authorize; it does not ship a prebuilt metric warehouse, and it does not write the ship decision back into production.

Implementation Steps after the Design Note

Start from the note, not from “is it significant.” Experiment analysis that starts at the badge will shop for a metric.

Replay assignment before outcomes

Confirm uniqueness, variant balance, and that outcomes sit after assignment. Experiment analysis on a join that uses first-event time instead of assigned_at will manufacture lift. Record sample-ratio mismatch. A 48/52 split on a 50/50 design is a diagnosis, not a footnote.

If you will peek, the note must already name the sequential rule. Daily refreshes are not a method. For horizon math, see A/B test sample size.

Bind guardrails and ask for the memo

List two or three kill metrics before you compute the primary. Typical experiment analysis kills: refund rate, p95 latency, support contacts, unsubscribe. If a join is missing, write “not measured.” Then ask for a decision memo: primary, interval, guardrails, CUPED on or off, balance, recommended action, and the SQL.

The agent writes the memo. A human decides ship, hold, or iterate. Open the query. Experiment analysis without attached SQL is a slide. If you need the downloadable artifact shape, continue in experiment decision memo. If you need the join itself, use analyze experiment results in SQL.

Desk Sample: Illustrative Assignment-First Pack

The following numbers are an illustrative desk composite, not a customer result and not an uplift claim.

ItemDesk composite (illustrative)
Window28 days after freeze
Units64,200 users; 50.4% / 49.6%
SRM noteMild 50.4/49.6; assignment job checked
PrimaryCheckout completion
Observed delta+1.1 percentage points (illustrative)
GuardrailRefund rate +0.3 percentage points (illustrative)
DecisionHold; refund join was late and still adverse

Experiment analysis on this pack is useful because assignment was read first. A memo that opened on +1.1 pp would have looked like a ship. The late refund join is the finding.

Grouped bar chart: Checkout Δ pp, Refund guardrail Δ pp, Assignment 50.4/49.6 × Read lift first vs Read assignment table first (desk composite from this page)

Figure. Desk composite from this page: 64,200 users; +1.1 pp checkout; +0.3 pp refund; mild SRM noted. Published context: census.gov; cms.gov; fda.gov. Not a customer experiment, SLA, or official benchmark.

Evidence classWhat you can citeWhat you cannot claim
Desk composite on this pageGrain, order, inspectable artifactsCustomer uplift or vendor bake-off
Published public sources aboveFrames, protocols, published extractsThat those agencies ran this desk pack

Desk composite: +1.1 pp primary, +0.3 pp refund → hold. Context: Census frames, CMS integrity, FDA protocol freeze, data.gov vintages, ONS exclusions.

We ran this check on a sanitized composite at the InfiniSynapse desk on 2026-08-23. We bound the note, then asked one experiment analysis question. We kept the memo only after the assignment table, the written primary, and both intervals were visible. We rejected lift before assignment. Figures stay illustrative. What you can copy is the assignment join and the on/off rule, not a lift.

Selection Scorecard for Post-Design Reads

Score from 1 to 5. Experiment analysis that cannot show assignment SQL should not win on a prettier badge.

CriterionWhat “5” looks likeDisqualifier
Assignment firstUnique unit, timestamps, SRMLift-first dashboards
Freeze honestyNote unchanged after peekPrimary rewritten mid-test
Inclusion namedITT or exposed, writtenSilent row drops
GuardrailsPre-registered and joined“Fine” with no column
AuditMemo + SQL downloadableChat-only winner text
Decision rightsNamed human ownerModel “recommends ship” as policy

Experiment analysis scores well when a skeptical partner can replay the read. It scores poorly when the stack implies a prebuilt experiment warehouse you do not operate.

Failure Modes That Ignore the Freeze

Write the break in the memo if it happened. Reviews go faster when invalidations are explicit.

Lift before assignment

Opening on conversion hides dual-exposed users and bad clocks. Experiment analysis that skips uniqueness will discover the bug in the launch email. Fix the log, then compute.

Thawing the primary after a peek

Changing the primary because the original looked flat is a new test. Experiment analysis may attach the exploratory chart. It may not treat the new metric as the frozen contract.

Guardrail silence after freeze

A conversion win that raises refunds is not a product win. Experiment analysis that cannot join a kill metric must say so. “Not in the extract” is a hold reason, not a pass.

A fourth pattern is treating the agent paragraph as the ship order. The memo is a draft. The owner’s name is the decision.

Before you open a workspace, check four things: unit uniqueness, assignment timestamps, a written primary, and at least one guardrail join. If those four are missing, a tool will still produce a confident interval.

Route the same diagnosis to the live guide that owns the next object. Each row is a single hop.

Live guideOpen it when
explainable AI data analysisthe plan and SQL must be auditable
ecommerce analyticsorders and refunds sit in different sources
data knowledge basemetric sentences need a bound note

Ask for the memo from the assignment table

Upload a sanitized assignment-and-outcome extract, bind the metric note, and ask for a memo that starts with assignment integrity. This check uses only sources you authorize.

Commercial association: You do not need the workspace to complete the educational diagnosis on this page.

Open InfiniSynapse

Use only authorized, sanitized data. Do not paste secrets.

How this page is sourced. William Zhu is cofounder of InfiniSynapse (GitHub @allwefantasy); no personal LinkedIn is published. Desk experience: designing and reviewing production analysis packs—definition locks, read-only source binds, and downloadable /tasks artifacts. Reviewed by analytics engineering · data platform · LLM security · editor. Editorial standards · corrections · publishing principles · Contact zhuhl@infinisynapse.com. Company Vision. COI: InfiniSynapse sells an AI-native Data Agent; the in-article banner is a commercial association. Fact-check: census.gov · cms.gov · fda.gov · data.gov · gov.uk.

Frequently Asked Questions

Do I read lift first if the design is already frozen?

Bottom line: No. Experiment analysis after freeze still starts with assignment. The freeze locks the contract; it does not prove the log is clean. Check uniqueness, timestamps, eligibility, and SRM before any lift sentence. If those fail, the interval is not a launch order.

Can I switch from ITT to exposed-only after seeing the number?

Bottom line: Not as the primary. Experiment analysis may report exposed-only as a secondary, labeled exploratory, if exposure is logged. Promoting it after a disappointing ITT is thawing the design. Keep both lines if you compute both, and say which one the owner is deciding on.

Is a CSV export enough for a frozen test?

Bottom line: Yes, if it includes unit, variant, assignment time, and outcomes. Experiment analysis on a file is still analysis. Freeze the file date. Bind metric definitions. Do not treat a missing refund column as a zero refund rate. If randomization lived elsewhere, keep those IDs.

Who signs the memo if an agent drafted it?

Bottom line: A human with launch authority. Experiment analysis can produce the intervals and the SQL. It cannot accept residual risk. Write the owner on the document. If marketing wants a public claim, raise the evidence bar before copy ships.

Conclusion

Experiment analysis after freeze is assignment first, then outcomes, then a memo. The freeze is a protocol: unit, primary, window, peeking rule. It is not a license to skip SRM or hide a missing guardrail. Hold when the log is dirty. Label slices that were not in the note. Name the human who ships.

When the assignment table and the metric note are ready, ask for that memo on an authorized extract at https://app.infinisynapse.com/. Open the SQL, keep the file, and reuse the same freeze on the next test.

Experiment Analysis after the Design Is Frozen (2026)