Incrementality testing: the quick answer增量测试:快速回答

Incrementality testing is a controlled experiment that estimates outcomes caused by a marketing treatment by comparing a treatment group with a valid untreated control or holdout group. The difference estimates what happened because of the activity, not merely what an attribution rule credited to it.

增量测试是一种受控实验,通过比较接受营销干预的处理组与有效的未处理对照组(留出组),估算由营销真正造成的结果。两组差异用于估计活动带来的新增结果,而不是重复平台归因规则分配的功劳。

The word incremental refers to the counterfactual difference: conversions, revenue, leads, visits, searches, or another predeclared outcome that would not have occurred in the same window without the treatment. A campaign may report many attributed conversions yet have modest incremental lift if it reaches people who were already likely to convert. The reverse is also possible when the platform's attribution window misses delayed or cross-device outcomes.

“增量”指反事实差异:如果没有这次干预,在相同窗口内原本不会发生的转化、收入、线索、访问、搜索或其他预先声明的结果。若活动触达的本就是高转化人群,平台可能报告大量归因转化,但真实增量有限;若归因窗口漏掉延迟或跨设备结果,也可能出现相反情况。

A valid test needs more than two lines on a chart. It needs a decision, eligible population, experimental unit, treatment, control condition, randomization or defensible geo construction, outcome definition, sample-size or power analysis, exposure safeguards, analysis plan, and uncertainty interval. If any of those changes after results appear, the estimate becomes easier to bias.

有效测试不只是图上两条线。它需要明确决策、符合条件的人群、实验单位、处理方式、对照条件、随机分配或可信的地域构造、结果定义、样本量或检验功效分析、曝光保护、分析计划与不确定性区间。若结果出现后再改变这些定义,估计就更容易受到偏差影响。

Why teams search for marketing incrementality为什么团队需要衡量营销增量

Attribution and incrementality answer different questions. Attribution observes touchpoints and applies rules or models to assign credit. Incrementality creates or approximates a credible world in which the treatment was absent, then compares outcomes. The first helps describe journeys and operate channels; the second helps decide whether an intervention creates enough additional value to justify its cost.

归因与增量回答不同问题。归因观察触点,并按规则或模型分配贡献;增量则构造或近似一个“没有接受干预”的可信世界,再比较结果。前者帮助描述旅程和运营渠道,后者帮助判断某项干预是否创造了足以覆盖成本的新增价值。

Budget decisions预算决策

Estimate incremental conversions, value, cost, and iROAS before scaling, reducing, or reallocating spend.

在扩量、缩减或重新分配预算前,估算增量转化、价值、成本与 iROAS。

Attribution calibration校准归因

Compare experimental lift with platform or multi-touch credit to identify systematic over- or under-crediting.

把实验增量与平台或多触点归因结果比较,识别系统性的过度或不足归因。

Targeting and suppression定向与排除

Test whether retargeting, discounts, or reminders move behavior beyond the baseline rather than harvest existing intent.

测试再营销、折扣或提醒是否改变基线行为,而不是只收割原本存在的意向。

Measurement triangulation多方法交叉验证

Use experiments as local causal anchors beside marketing mix modeling, attribution, forecasting, and qualitative evidence.

将实验作为局部因果锚点,与营销组合模型、归因、预测和定性证据共同使用。

Do not run a test merely because a dashboard metric looks suspicious. State the action that would change at plausible outcomes. If no realistic result would alter a campaign, budget, audience, or measurement policy, the opportunity cost of withholding treatment may outweigh the information value.

不要只因为仪表板指标可疑就启动实验。应先说明在不同合理结果下,哪些行动会改变。如果没有任何现实结果会影响活动、预算、人群或衡量政策,那么留出干预所造成的机会成本可能高于信息价值。

When an incrementality test is—and is not—the right method何时适合做增量测试,何时不适合

Decision决策Preferred method优先方法Why原因
Did this campaign or channel cause additional conversions?该活动或渠道是否造成新增转化?User holdout or geo incrementality test用户留出或地域增量实验Directly contrasts treatment with a no-treatment counterfactual直接比较处理与未处理反事实
Which headline, offer, or landing page performs better?哪个标题、优惠或落地页表现更好?A/B or multivariate testA/B 或多变量测试Compares executions; include a no-treatment arm only if total tactic lift matters比较执行版本;仅当需要判断整体策略增量时加入无处理组
How did many channels contribute over several years?多个渠道多年间如何贡献?Marketing mix modeling plus experiments营销组合模型结合实验MMM covers aggregate history; experiments can calibrate selected effectsMMM 覆盖汇总历史,实验校准部分效应
Which touchpoints appeared before a conversion?转化前出现了哪些触点?Journey analysis or attribution旅程分析或归因Describes observed paths but does not by itself identify a counterfactual描述已观察路径,但本身不能识别反事实
What is the long-run brand or market equilibrium effect?长期品牌或市场均衡效应是什么?Repeated studies, surveys, MMM, and strategic research重复实验、调查、MMM 与战略研究A short holdout can miss delayed, spillover, and system-level effects短期留出会漏掉延迟、溢出和系统层效应

User-level randomization is usually strongest when individuals can be assigned before treatment, identities are stable enough for outcome measurement, and spillovers are limited. Geo experiments are useful when user suppression is unavailable, cross-device identity is weak, offline outcomes matter, or a channel is purchased geographically. Geo designs trade more assumptions and fewer units for broader outcome coverage.

当用户可在处理前分配、身份足以稳定衡量结果且溢出有限时,用户级随机实验通常最有力。若无法排除单个用户、跨设备身份较弱、线下结果重要,或渠道按地域购买,可考虑地域实验。地域设计以更多假设和更少实验单位,换取更广的结果覆盖。

Do not test when treatment cannot be withheld safely or ethically. Also pause when outcome volume is too low, groups will contaminate each other, the campaign changes simultaneously across markets, or the decision window is shorter than the conversion lag. Use a pilot, broader outcome, longer window, aggregated design, or observational analysis with clearly weaker causal language.

若无法安全或合乎伦理地停止干预,就不要做留出测试。当结果量过低、组间必然污染、各市场同时发生重大变化,或决策窗口短于转化延迟时也应暂停。可改用试点、更上层结果、更长窗口、汇总设计,或采用明确降低因果表述强度的观察性分析。

Data and prerequisites for a defensible lift test可信提升测试所需的数据与前提

Begin with a one-page decision brief, not a query. Record the owner, treatment, eligible population, experimental unit, primary outcome, acceptable risk, minimum effect worth acting on, planned duration, exclusions, contamination threats, and actions for positive, negative, and inconclusive results. Freeze this brief before looking at test outcomes.

从一页决策简报开始,而不是先写查询。记录负责人、干预、符合条件的人群、实验单位、主要结果、可接受风险、值得采取行动的最小效应、计划时长、排除规则、污染风险,以及正向、负向和不确定结果分别对应的行动。查看测试结果前冻结这份简报。

  • Eligibility and identity: stable user, household, store, or geography keys; eligibility timestamp; consent or suppression rules; mutually exclusive assignment.
  • 资格与身份:稳定的用户、家庭、门店或地域键;资格时间;同意或排除规则;互斥分组。
  • Assignment: treatment/control flag, assignment time, randomization stratum or matched-market set, intended treatment, and any reassignment.
  • 分组:处理/对照标记、分配时间、随机分层或匹配市场集合、计划干预及任何重新分配。
  • Delivery: impressions, sends, bids, availability, spend, treatment start and stop, failed deliveries, and evidence that holdouts stayed untreated.
  • 交付:曝光、发送、出价、可投放状态、花费、处理起止、失败交付,以及留出组确实未处理的证据。
  • Outcomes: event name, timestamp, value, currency, deduplication key, attribution-independent collection rule, refund or cancellation policy, and conversion lag.
  • 结果:事件名称、时间、价值、币种、去重键、与归因无关的采集规则、退款或取消政策及转化延迟。
  • Context: pre-period outcomes, seasonality, promotions, stock, outages, competitor events, other channel changes, and geographic migration where relevant.
  • 背景:实验前结果、季节性、促销、库存、故障、竞争事件、其他渠道变化,以及适用时的地域迁移。

Reconcile totals before analysis. Treatment plus control assignment should equal the eligible population after documented exclusions. Outcome events should reconcile to a trusted order, CRM, or analytics total under the same date and status rules. Cost should match finance or platform spend at the experimental grain. A statistically elegant model cannot repair silent population loss or mismatched time zones.

分析前先对账。按记录的排除规则处理后,处理组与对照组分配总数应等于符合条件的人群。结果事件应在相同日期与状态规则下,与可信订单、CRM 或分析总数一致;成本也应在实验粒度上与财务或平台花费一致。再精巧的统计模型也无法修复悄然丢失的人群或错位的时区。

Choose the right incrementality testing design选择合适的增量测试设计

Design设计Best fit适用场景Main risk主要风险Validation验证重点
User-level randomized holdout用户级随机留出Addressable ads, email, product messages, offers可寻址广告、邮件、产品消息、优惠Identity loss, cross-device exposure, noncompliance身份丢失、跨设备曝光、不遵从分配Assignment balance, delivery leakage, intent-to-treat分配平衡、交付泄漏、意向处理分析
Geo-randomized or matched-market test地域随机或匹配市场测试Offline media, broad digital channels, store or regional outcomes线下媒体、广覆盖数字渠道、门店或区域结果Few units, market shocks, spillover, weak matches单位少、市场冲击、溢出、匹配较弱Pre-fit, placebo tests, sensitivity to market selection实验前拟合、安慰剂测试、市场选择敏感度
Switchback or time-block test切换式或时间区块测试Marketplace, delivery, pricing, or systems that cannot split users无法按用户分组的市场、配送、定价或系统Carryover, trends, daypart imbalance延续效应、趋势、时段不平衡Washout, randomized blocks, robust time controls清洗期、随机区块、稳健时间控制
PSA or placebo control公益广告或安慰剂对照Platforms where an untreated impression opportunity must be observed需要观察未处理曝光机会的平台Placebo may itself change behavior; cost and auction effects安慰剂本身影响行为;成本和拍卖效应Neutral creative, equal eligibility, delivery comparability中性素材、相同资格、交付可比性

Randomize at the level where interference is acceptably low. If household members influence each other, randomize households. If media can only be withheld by region, randomize or match regions. If one store's pricing changes demand at a nearby store, consider clusters. Standard errors and effective sample size must reflect that unit; thousands of users inside ten geographies do not create thousands of independent experimental units.

应在干扰可接受的层级随机分配。若家庭成员互相影响,就按家庭分配;若媒体只能按区域排除,就随机或匹配区域;若一家门店定价会改变附近门店需求,则考虑集群。标准误与有效样本量必须对应这一单位;十个地域中有数千用户,并不等于有数千个独立实验单位。

Power analysis should use the baseline rate, outcome variance, allocation ratio, clustering, decision-relevant minimum detectable effect, significance threshold, desired power, expected attrition, and multiplicity. Do not choose a holdout size from a generic percentage. Smaller holdouts reduce opportunity cost but may require more time or a higher-volume outcome; larger holdouts improve precision but can forgo more treatment value.

功效分析应纳入基线率、结果方差、分配比例、聚类、对决策有意义的最小可检测效应、显著性阈值、目标功效、预期流失和多重比较。不要套用通用留出比例。较小留出组降低机会成本,但可能需要更长时间或更高频结果;较大留出组提高精度,却可能放弃更多干预价值。

How to run an incrementality test step by step如何逐步执行增量测试

  1. Write the decision and estimand写清决策与估计目标Name the intervention, population, outcome, follow-up window, unit, and effect you need: absolute lift, relative lift, incremental value, or iROAS. Predeclare what result changes the decision.明确干预、人群、结果、跟踪窗口、单位及所需效应:绝对提升、相对提升、增量价值或 iROAS。预先声明什么结果会改变决策。
  2. Check feasibility and power检查可行性与功效Estimate baseline volume and variance from a representative pre-period. Model realistic lift, holdout opportunity cost, clustering, conversion lag, and attrition. Stop if the study cannot resolve a useful effect.用代表性实验前窗口估计基线量与方差,建模现实提升、留出机会成本、聚类、转化延迟和流失。若研究无法分辨有用效应,应停止。
  3. Create and freeze assignment创建并冻结分组Randomize with a reproducible seed and stratification, or document matched-market selection. Store assignment before treatment. Do not move poor performers between groups after launch.使用可复现随机种子与分层,或记录匹配市场选择。处理前保存分组。上线后不得因表现不佳而移动单位。
  4. Validate the launch验证上线执行Check eligible counts, sample ratio, pre-treatment balance, delivery, holdout suppression, time zones, spend, and concurrent campaigns. Diagnose implementation without peeking at treatment effect.检查资格人数、样本比例、处理前平衡、交付、留出排除、时区、花费和并行活动。在不偷看处理效应的情况下诊断执行。
  5. Run for the declared window按预设窗口运行Protect assignment and campaign settings. Log outages, promotions, supply constraints, policy changes, and cross-channel shocks. Avoid optional stopping because an early chart looks favorable.保护分组和活动设置,记录故障、促销、供应限制、政策变化与跨渠道冲击。不要因早期图表有利而选择性提前停止。
  6. Estimate effect and uncertainty估计效应与不确定性Analyze all assigned units under intent-to-treat by default, calculate the predeclared metrics, use methods consistent with randomization and clustering, and report confidence or credible intervals beside point estimates.默认按意向处理分析全部已分配单位,计算预先声明的指标,使用与随机化和聚类一致的方法,并在点估计旁报告置信区间或可信区间。
  7. Stress-test and decide压力测试并决策Check contamination, missingness, balance, placebo periods, alternative reasonable windows, extreme markets, refunds, and cost definitions. Then map the validated range—not only the point estimate—to the decision rule.检查污染、缺失、平衡、安慰剂时期、其他合理窗口、极端市场、退款与成本定义。最后将经验证的范围而非单一点估计映射到决策规则。

A monitoring dashboard may display delivery and data quality during the test, but the effect estimate should follow the stopping rule. If safety or operational harm requires early termination, stop the treatment and label the analysis accordingly; preserving business or participant safety matters more than statistical convenience.

测试期间可以用监控仪表板查看交付与数据质量,但效应估计应遵守停止规则。若安全或运营损害要求提前终止,应停止干预并如实标注分析;保障业务或参与者安全比统计便利更重要。

How to calculate incremental lift, conversions, and iROAS如何计算增量提升、增量转化与 iROAS

For equally followed user-level groups, let the treatment conversion rate be RT and the control rate be RC. Absolute lift is the percentage-point difference. Relative lift scales that difference by the control baseline. To estimate incremental conversions in a treatment group of size NT, multiply absolute lift by its eligible unit count. When exposure compliance is imperfect, this intent-to-treat estimate remains tied to the policy of assigning treatment.

对于跟踪条件相同的用户级组,设处理组转化率为 RT,对照组为 RC。绝对提升是百分点差,相关提升则用对照基线缩放这一差异。要估算规模为 NT 的处理组增量转化,可将绝对提升乘以符合条件的处理单位数。当曝光遵从不完全时,这一意向处理估计仍对应“分配处理”政策。

Absolute lift = RT − RC
Relative lift = (RT − RC) ÷ RC
Incremental conversions = (RT − RC) × NT
iROAS = Incremental conversion value ÷ Incremental cost

Hypothetical example: an eligible audience is randomized 80/20. The treatment has 80,000 users and 3,360 conversions, a 4.20% rate. The holdout has 20,000 users and 760 conversions, a 3.80% rate. Absolute lift is 0.40 percentage points; relative lift is about 10.5%; estimated incremental conversions among treated users are 320. If the predeclared net value per conversion is $70, incremental value is $22,400. If incremental campaign cost versus control is $16,000, point-estimate iROAS is 1.40.

假设示例:符合条件的人群按 80/20 随机分配。处理组 80,000 人、3,360 次转化,转化率 4.20%;留出组 20,000 人、760 次转化,转化率 3.80%。绝对提升为 0.40 个百分点,相对提升约 10.5%,处理组估算增量转化为 320。若预先声明每次转化净价值为 70 美元,则增量价值为 22,400 美元;若相对对照的增量活动成本为 16,000 美元,则 iROAS 点估计为 1.40。

Those arithmetic results are not yet a decision. Calculate an interval using a method appropriate to the outcome and assignment. If the plausible range for iROAS crosses the decision threshold, the evidence may be inconclusive even when the point estimate looks attractive. Also test sensitivity to value definition, refunds, conversion lag, missing outcomes, and clustered assignment. Report absolute and relative lift together; relative percentages can make small baselines look dramatic.

这些算术结果尚不足以决策。应使用适合结果类型和分配方式的方法计算区间。若 iROAS 的合理范围跨越决策阈值,即便点估计诱人,证据仍可能不确定。还应检验价值定义、退款、转化延迟、缺失结果和集群分配的敏感度。绝对与相对提升应同时报告,因为较小基线会让相对百分比显得夸张。

Validate an incrementality test before acting采取行动前验证增量测试

Validation should separate design, execution, data, analysis, and decision. Passing one layer does not compensate for failure elsewhere. A randomized assignment with contaminated holdouts is not a clean experiment; a significant estimate from unreconciled outcomes is not trustworthy; and a precise short-term result may still be irrelevant to a long-term decision.

验证应区分设计、执行、数据、分析和决策。某一层通过不能弥补其他层失败。随机分配但留出组受污染,不是干净实验;未对账结果得到的显著估计不可信;精确的短期结果也未必适用于长期决策。

  • Assignment: sample ratio matches plan; stratification and cluster rules are preserved; pre-treatment covariates and outcomes are balanced within expected random variation.
  • 分配:样本比例符合计划;分层和集群规则保持;处理前协变量与结果在预期随机波动内平衡。
  • Delivery: treatment eligibility, spend, reach, frequency, and timing match the intervention; control leakage and cross-channel substitution are quantified.
  • 交付:处理资格、花费、覆盖、频次和时间与干预一致;量化对照泄漏与跨渠道替代。
  • Outcome data: collection is assignment-blind where possible; duplicate, refund, cancellation, currency, time-zone, and late-arrival rules reconcile to trusted totals.
  • 结果数据:尽可能在不知道分组的情况下采集;去重、退款、取消、币种、时区和延迟到达规则与可信总数对账。
  • Inference: estimator follows the pre-analysis plan; standard errors reflect the randomization unit; multiplicity, missingness, and optional stopping are handled.
  • 推断:估计器遵循预分析计划;标准误反映随机化单位;妥善处理多重比较、缺失和选择性停止。
  • Robustness: conclusions survive reasonable window, outlier, covariate-adjustment, placebo, and market-selection checks without cherry-picking.
  • 稳健性:结论能通过合理窗口、异常值、协变量调整、安慰剂和市场选择检查,而非挑选有利结果。
  • Decision fit: interval, external validity, duration, treatment intensity, and cost definition match the policy or budget move under consideration.
  • 决策匹配:区间、外部有效性、持续时间、处理强度和成本定义与待定政策或预算调整匹配。

Treat a negative point estimate carefully. It may indicate harm, random noise, displaced conversions, supply constraints, poor execution, or a control group exposed elsewhere. Treat “no lift” carefully too: failure to reject zero is not evidence of equivalence unless the test was designed and analyzed for an equivalence margin. Publish the interval and minimum detectable effect so readers can see what the study could and could not rule out.

应谨慎解读负点估计:它可能代表伤害、随机噪声、转化转移、供应限制、执行不佳或对照组在其他位置受曝光。“无提升”也需谨慎:除非研究围绕等效界值设计并分析,否则未能拒绝零效应并不等于证明等效。应发布区间和最小可检测效应,让读者了解研究能排除和不能排除什么。

Common incrementality testing mistakes, limits, and risks增量测试的常见错误、局限与风险

Choosing groups after the result结果出现后再选组

Post-treatment matching or removing weak markets can manufacture lift. Freeze eligibility, assignment, exclusions, and market selection first.

处理后匹配或移除弱市场会制造提升。应先冻结资格、分组、排除与市场选择。

Measuring exposure instead of assignment按曝光而非分组分析

Exposed users are often selected by auctions or behavior. Default to intent-to-treat; estimate treatment-on-treated only with justified methods and assumptions.

曝光用户常受拍卖或行为选择。默认意向处理;仅在方法与假设合理时估计实际接受处理者效应。

Calling the test too early过早停止测试

Repeated peeking inflates false-positive risk under fixed-horizon analysis. Follow the declared horizon or use a valid sequential design.

固定期限分析中反复偷看会提高假阳性风险。应遵循预设周期或采用有效序贯设计。

Ignoring interference忽视干扰

People share offers; media crosses borders; inventory and auctions react. Randomize clusters, create buffers, or narrow the causal claim.

人会分享优惠,媒体跨区域传播,库存与拍卖会反应。可按集群分配、设置缓冲区或收窄因果主张。

Confusing short-term with total value把短期效应当成总价值

A test window may miss repeat purchase, brand effects, churn, pull-forward, and channel interactions. State the measured horizon explicitly.

测试窗口可能漏掉复购、品牌效应、流失、需求前置和渠道互动。必须明确说明衡量周期。

Using platform metrics as ground truth把平台指标当作唯一真相

Platform studies can be useful, but eligibility, modeling, unavailable accounts, attribution-independent outcomes, and reporting rules still require review.

平台研究有用,但仍需检查资格、建模、账户可用性、独立结果采集和报告规则。

External validity is a separate question. An estimate applies most directly to the tested population, treatment intensity, creative, season, auction environment, geography, and outcome window. Scaling can change reach, frequency, audience quality, price, or competitive response. Repeat tests when those mechanisms change materially, and maintain a registry so selective publication does not turn one favorable study into policy.

外部有效性是另一问题。估计最直接适用于被测试的人群、处理强度、素材、季节、拍卖环境、地域与结果窗口。扩量会改变覆盖、频次、人群质量、价格或竞争反应。若机制发生实质变化,应重复测试,并维护实验登记,避免选择性发布把一次有利结果变成长期政策。

Analyze prepared incrementality test data with InfiniSynapse使用 InfiniSynapse 分析已准备的增量测试数据

Before opening a tool, prepare an assignment table, delivery or exposure log, outcome table, spend table, metric dictionary, decision brief, and exclusion rules. Reconcile the eligible population and totals. Keep raw assignment immutable, and identify the randomization unit and time zone in plain language.

打开工具前,请准备分组表、交付或曝光日志、结果表、花费表、指标字典、决策简报与排除规则。完成符合条件人群和总数对账,保持原始分组不可变,并用清晰语言标注随机化单位与时区。

Turn prepared experiment exports into a reviewable analysis把已准备的实验导出转化为可复核分析

Use the InfiniSynapse AI Data Analyst to inspect files and connected databases, reconcile treatment and control counts, calculate declared metrics, explore diagnostics, and produce tables or summaries for analyst review. InfiniSynapse is not described here as an ad-delivery system, randomization service, power calculator, or automatic proof of causality.

可使用 InfiniSynapse AI Data Analyst 检查文件与已连接数据库、对账处理与对照人数、计算预设指标、探索诊断,并生成供分析师复核的表格或摘要。本页不会把 InfiniSynapse 描述为广告投放系统、随机分组服务、功效计算器或自动因果证明工具。

Analyze prepared experiment data分析已准备的实验数据

A bounded first request might be: “Using assignment as the source of group membership, reconcile eligible counts; calculate treatment and control outcome rates for the declared 28-day window; report absolute lift, relative lift, incremental outcomes, incremental value, and iROAS; preserve cluster IDs; flag contamination and missing outcomes; and show formulas plus uncertainty.” An analyst should still verify the estimator, standard errors, power, assumptions, and causal interpretation.

一个边界清晰的首个请求可以是:“以分组表作为组别唯一来源,对账符合条件人数;按预设 28 天窗口计算处理与对照结果率;报告绝对提升、相对提升、增量结果、增量价值和 iROAS;保留集群 ID;标记污染与缺失结果;展示公式与不确定性。”分析师仍需验证估计器、标准误、功效、假设与因果解释。

For broader multi-channel historical allocation, continue with the marketing mix modeling guide. For recurring joins across ad, web, CRM, and warehouse data, use the marketing analytics guide. These adjacent workflows support planning and operations questions; neither replaces a well-designed experiment.

若要进行更广泛的跨渠道历史预算分配,请继续阅读营销组合模型指南。若要持续连接广告、网站、CRM 与数仓数据,请参考营销分析指南。这些相邻工作流支持规划与运营问题,但都不能替代良好设计的实验。

Incrementality testing best practices and next steps增量测试最佳实践与下一步

  • Pre-register the decision, not only the metric. Write what you will increase, decrease, stop, repeat, or investigate across plausible intervals.
  • 预登记决策,而不只是指标。写清在不同合理区间下会增加、减少、停止、重复或调查什么。
  • Use one primary outcome. Secondary outcomes can explain mechanisms, but a large menu invites selective conclusions and multiplicity.
  • 只设一个主要结果。次要结果可以解释机制,但过多指标会诱发选择性结论和多重比较。
  • Design around the decision threshold. Power the study for the smallest effect that would change action, not an optimistic benchmark.
  • 围绕决策阈值设计。以足以改变行动的最小效应进行功效设计,而不是乐观基准。
  • Automate quality checks before effect estimation. Assignment counts, sample ratio, missingness, exposure leakage, spend, and outcome reconciliation should fail loudly.
  • 在效应估计前自动化质量检查。分组人数、样本比例、缺失、曝光泄漏、花费和结果对账出现问题时必须明显报错。
  • Report intervals and operational context. Pair every lift estimate with uncertainty, unit, window, treatment intensity, cost definition, and known deviations.
  • 报告区间与运营背景。每个提升估计都应同时给出不确定性、单位、窗口、处理强度、成本定义和已知偏差。
  • Build a measurement portfolio. Use experiments for local causal questions, MMM for aggregate history and scenarios, attribution for journeys and operations, and qualitative research for mechanisms.
  • 构建衡量组合。用实验回答局部因果问题,用 MMM 研究汇总历史与情景,用归因处理旅程和运营,用定性研究解释机制。

The practical next step is a feasibility analysis. Choose one decision with meaningful spend, stable delivery, enough outcome volume, and a control condition you can maintain. Pull a representative pre-period, define a minimum useful effect, simulate power and opportunity cost, list spillover paths, and review the plan with channel, data, finance, privacy, and operations owners. Launch only after each owner understands what will be withheld and how the result will be used.

务实的下一步是做可行性分析。选择一个花费重要、交付稳定、结果量充足且能维持对照条件的决策;提取代表性实验前窗口,定义最小有用效应,模拟功效与机会成本,列出溢出路径,并与渠道、数据、财务、隐私和运营负责人共同审核。只有所有负责人都理解将排除什么、结果如何使用后,才应上线。

Frequently asked questions about incrementality testing关于增量测试的常见问题

What is incrementality testing?什么是增量测试?

Incrementality testing is a controlled experiment that estimates outcomes caused by a marketing treatment by comparing a treatment group with a valid untreated control or holdout group.

增量测试是一种受控实验,通过比较处理组与有效的未处理对照组或留出组,估算由营销干预真正造成的结果。

How does incrementality testing work?增量测试如何工作?

Define one decision and outcome, assign comparable units to treatment and holdout conditions, preserve the assignment, measure both groups over a predeclared window, estimate the difference with uncertainty, and check for contamination or operational failures.

先定义一个决策和结果,将可比单位分配到处理与留出条件,保持分组不变,在预设窗口内衡量两组,估计带不确定性的差异,并检查污染或运营失败。

Is incrementality testing the same as A/B testing?增量测试与 A/B 测试相同吗?

They share experimental principles, but they often answer different questions. A creative A/B test asks which version performs better; an incrementality test usually asks whether receiving the activity creates outcomes beyond receiving none of it.

两者共享实验原则,但常回答不同问题。素材 A/B 测试比较哪个版本更好;增量测试通常判断接受活动是否比完全不接受活动创造更多结果。

What data do you need for an incrementality test?增量测试需要哪些数据?

You need stable unit identifiers, assignment and eligibility records, treatment delivery or exposure evidence, a predeclared outcome with event time and value, cost, exclusion flags, and enough pre-period data to diagnose balance or match geographies.

需要稳定的单位标识、分组与资格记录、干预交付或曝光证据、带事件时间与价值的预设结果、成本、排除标记,以及足够的实验前数据用于诊断平衡或匹配地域。

What if an incrementality test is inconclusive?如果增量测试结论不明确怎么办?

An inconclusive result means the study did not estimate the effect precisely enough for the chosen decision; it does not prove zero effect. Report the interval, diagnose power and execution, and redesign only when another test can materially reduce uncertainty.

不明确表示研究对当前决策所需效应的估计不够精确,并不证明效应为零。应报告区间、诊断功效与执行,并仅在新测试能实质降低不确定性时重新设计。

Can incrementality testing measure long-term effects?增量测试能衡量长期效应吗?

It can measure outcomes inside an ethically and operationally feasible follow-up window. Delayed, spillover, brand, or equilibrium effects may require longer follow-up, repeated studies, surveys, geo designs, or triangulation with marketing mix modeling.

它能衡量伦理和运营上可行的跟踪窗口内结果。延迟、溢出、品牌或均衡效应可能需要更长跟踪、重复研究、调查、地域设计,或与营销组合模型交叉验证。

Official sources for lift studies and geo experiments提升研究与地域实验的官方来源