B has the higher signup rate. Why does it lose on both devices?
B converts 25% of users; A converts 23%. That looks like a rollout decision until desktop and mobile each point the other way. We gave ChatGPT, Gemini and Claude this fictional signup test. WekeyLab AI would use ChatGPT’s decision note as the starting point: its calculations are right, and it keeps the rollout on hold while asking how users reached each design.
Start with ChatGPT’s decision, then shorten the note
All three calculated the four conversion rates correctly. The difference was what they told the team to do with those rates. ChatGPT separated “B had a higher observed rate” from “B caused an improvement.” That distinction is the reason to choose its note here.
Gemini explained the reversal clearly, then predicted that shipping B would probably reduce signups. Claude added useful uncertainty calculations, but its assignment verdict and rerun plan need more conditional wording. I would take the checks from both, keep ChatGPT’s decision boundary, and put the team’s next tasks in one memo.
The four rows the team needs to see
Both designs enrolled users during the same 14 days. Every user has a full seven-day follow-up, appears in one design and one device group, and can count as a signup once. Device means device at first exposure. The export is complete; assignment, targeting and allocation changes have not yet been checked.
Here is the complete input table. Conversion is completed signups divided by exposed users. Desktop and mobile are the only device groups.
| Design | Device | Exposed | Signups | Rate |
|---|---|---|---|---|
| A | Desktop | 300 | 90 | 30% |
| A | Mobile | 700 | 140 | 20% |
| B | Desktop | 700 | 196 | 28% |
| B | Mobile | 300 | 54 | 18% |
Show input CSV
Design,Device,Exposed users,Completed signups A,Desktop,300,90 A,Mobile,700,140 B,Desktop,700,196 B,Mobile,300,54
Put the total beside both device groups
Observed conversion · A is filled, B is outlined
Overall
B is 2 pp higher
Desktop
B is 2 pp lower
Mobile
B is 2 pp lower
All charts use the same 0–40% scale
What happens with the same device mix?
Desktop users were 30% of A and 70% of B. Give both designs the same desktop weight. In this example, B remains 2 percentage points lower at every common weight.
At a 50/50 desktop/mobile mix, A is 25% and B is 23%. B − A is −2 percentage points.
Try a different device mix
Mobile weight: 50%
Recalculated at a common mix
This recombines observed device rates. Check assignment and uncertainty before interpreting a design effect.
The 8.7% is calculated correctly. “Improves” is the problem
A has 90 + 140 = 230 signups out of 1,000 users: 23%. B has 196 + 54 = 250 out of 1,000: 25%. Subtract the rates and the difference is 2 percentage points. Divide that difference by A’s 23% and the relative difference is about 8.7%. It is not an 8.7 percentage-point increase.
Now split the same data by device. Desktop is A 30% versus B 28%. Mobile is A 20% versus B 18%. B is 2 percentage points lower in each group. The combined result points the other way because B’s users are 70% desktop, while A’s are only 30% desktop. Desktop has the higher observed rate under both designs.
For a transparent comparison, give both sets of observed rates a 50/50 device mix: A becomes 0.5×30% + 0.5×20% = 25%; B becomes 0.5×28% + 0.5×18% = 23%. These are reweighted summaries, not a forecast of a rollout. The reversal is commonly called Simpson’s paradox: groups and their combined total point in opposite directions.
ChatGPT: the rollout decision stays attached to the missing evidence
ChatGPT calculates both the absolute and relative differences, then demonstrates the mix problem using A’s weights and B’s weights separately. Its strongest sentence is about assignment, not a winner. It asks for the randomization method, targeting rules, changes during the 14 days and uncertainty estimates.
It also asks whether the primary outcome was chosen before the results were seen. That is worth keeping in the handoff. The answer is repetitive, though: its analysis and decision note both walk through the same counts, rates and device mix. I would retain the calculation once and move directly to owners and required records.
ChatGPT: “First establish how users were assigned.”
Gemini: holding B is sensible; predicting a decline goes too far
Gemini’s explanation is the shortest of the three and quickly identifies the device-mix reversal. It also asks about allocation changes and downstream activation or retention. Those are useful checks.
Its next step says shipping B “would likely decrease your overall conversion rate.” Lower observed rates in both groups do not establish that prediction while assignment and statistical uncertainty remain unresolved. The action can stay firm—hold the rollout—without turning A into a proven winner.
A second sentence needs replacing: it says to continue until significance is reached in each segment. A difference may never become significant. Set a stopping rule and maximum horizon before collecting more data. Continuous checking needs a method designed for repeated looks; simply repeating a conventional fixed-sample test is a different procedure. Amplitude’s statistical-settings documentation distinguishes those methods.
Gemini: “Re-run or continue the test with balanced allocation until statistical significance is reached within each device segment.”
Claude: useful ranges, but put the assumptions beside the numbers
Claude supplies approximate 95% confidence intervals. Recalculating its normal approximations gives about −1.7 to +5.7 percentage points for the pooled B-minus-A difference, and −6.1 to +2.1 points after a fixed 50/50 device weighting. Both contain zero. The arithmetic supports its refusal to declare B worse. Interpreting these as design effects still depends on valid assignment and sampling assumptions.
It also calculates a device-by-design chi-square statistic of 320, which checks out against the independence expectation of 500 users per cell. That is a strong reason to investigate the split. I would replace its categorical “Assignment was not independent of device” with a request to compare the observed split with the intended allocation and logs. We have not inspected the assignment mechanism.
The approximate 7,000 users per arm is plausible for an illustrative equal-arm, fixed-horizon test detecting a change from 23% to 25%, with 80% power and a two-sided 5% threshold; a standard calculation gives about 7,157. A device-specific design or a different target effect needs its own calculation. Likewise, “one more test cycle” is a rerun proposal, not a known cost of delay.
Work through the rollout decision
Hold the all-user rollout of B. Audit assignment, then choose between analysing usable existing data and designing a new experiment.
Read the steps
- Hold the full rollout of B and audit allocation, changes and device exposure.
- If the assignment conditions support comparison, analyse existing data with explicit assumptions and uncertainty. Otherwise design a new experiment, specifying allocation, metrics, sample size and a stopping rule.
- Product, experiment and data owners review the evidence and open questions.
- If evidence is insufficient, resume the relevant check or experiment and review again. Otherwise record the rollout decision and conditions.
What I would ask the team to leave behind
First, the experiment owner exports the assignment rules, intended split and configuration change history. The data analyst joins daily assignment and exposure counts by design and device, then checks whether campaigns, countries or time periods differ too. The outcome is an assignment audit, not just another conversion chart.
Next, the analyst records which existing observations support a defensible comparison and which assumptions are needed. If the setup can support analysis, choose the relevant population mix and quantify uncertainty. If it cannot, prepare a new test plan with assignment, the primary outcome, guardrails, target effect, sample size and stopping rule specified together. The product owner decides between these paths after the audit.
I would keep three records linked by descriptive names: the assignment audit, the calculation sheet and the rollout decision. Each open item gets an owner and an agreed review date. That lets the next person see which missing fact prevented a launch and what evidence would change the decision.
WekeyLab AI’s decision note
Copy this note with the four-row example. For another experiment, replace the counts, time windows, groups and open items together.
Signup-page decision: hold the all-user rollout of B. Observed result: A 230/1,000 = 23%; B 250/1,000 = 25%. B is 2 percentage points higher overall, or about 8.7% relatively. Device check: desktop A30% / B28%; mobile A20% / B18%. Desktop accounts for 30% of A users and 70% of B users. The overall gap does not establish a design improvement. Decision: do not publish “B improves conversion by 8.7%” as a causal claim. Neither design is established as the causal winner from this export. Before reconsidering: experiment owner to provide assignment rules, intended split and change history; analyst to check daily exposure, targeting differences and an appropriate uncertainty analysis. Owners and review date: [agree with the team]. Next gate: review that audit, then decide whether to analyze valid existing data or design a fresh test. Record the chosen population, primary metric, guardrails and stopping rule. Suggested interim wording: “B’s observed signup rate was 25% versus A’s 23%. Device mix differed substantially, so the rollout decision is on hold pending assignment and analysis checks.”
How I would use AI on the next readout
Give it the unit counted, the exposure event, the conversion window, exclusions and raw counts before asking for a decision. Separate known setup facts from questions still waiting for answers. The full prompt used here and all three originals are available below.
My next instruction would be: “Keep observed results separate from causal claims. Recalculate totals and subgroup rates, explain any change in direction, and list the records that would let us decide. Put owners and undecided items into the memo without inventing agreement.” This is a reusable follow-up instruction, not an additional response collected for this comparison.
The related survey-evidence comparison covers how a sample supports a conclusion; the project-status comparison covers how to carry unknowns into a team report. Here the useful habit is the same: write the number, its meaning and the action it actually supports.
Sources and actual records
Collected September 24, 2026, 09:04–09:12 KST from the actual websites: ChatGPT Free, Think off, specific model not displayed; Gemini Flash, plan unverified; Claude Free, Sonnet 5 Medium. One identical English prompt per service, no length cap or answer-quality retries. Fictional input and eight criteria fixed before submission; original-copy fingerprints and 24 criterion reviews preserved. Additional normal approximations independently recalculated under stated assumptions. Korean and Chinese pages localize this same English run. Raw Markdown lengths including formatting and line breaks: ChatGPT 5,746 characters, Gemini 4,071, Claude 5,896. Search volume was not established; no real rollout was performed.
Google Analytics — A/B testing
Amplitude — Statistical methods and stopping rules
Amplitude — Troubleshoot experiment assignment and exposure
ChatGPT · 2026-09-24T00:07:04.355572+00:00
I'm a product manager reviewing two signup-page designs. Our team wants to ship design B to everyone because the dashboard says its signup conversion rate is higher. Can you analyze the results, tell me what you would do next, and draft the decision note I can take to the team? This is a fictional exercise. The export covers the same 14-day enrollment period for both designs, and every included user has a complete seven-day follow-up. Each row counts distinct eligible users who first saw that design, with completed signups counted once within seven days of that exposure. A user appears in only one design and one device group; device is recorded at first exposure. Desktop and mobile are the only groups. There are no missing counts in this export. I have not yet checked how users were assigned to designs, whether allocation changed during the period, or how targeting was configured. Design A, desktop: 300 exposed users, 90 completed signups. Design A, mobile: 700 exposed users, 140 completed signups. Design B, desktop: 700 exposed users, 196 completed signups. Design B, mobile: 300 exposed users, 54 completed signups. The proposed announcement says: "B improves signup conversion by 8.7%, so we should roll it out to all users." I need to know whether that claim and rollout decision follow from these numbers. Please show the calculations that matter and explain what evidence we still need. The raw numbers do show a higher **overall observed conversion rate** for B, but they do **not by themselves establish that the design caused the improvement** or justify an immediate all-user rollout. ### 1. What the export actually shows | Design | Desktop | Mobile | Overall | | ------ | --------------------: | --------------------: | ----------------------: | | A | 90 / 300 = **30.0%** | 140 / 700 = **20.0%** | 230 / 1,000 = **23.0%** | | B | 196 / 700 = **28.0%** | 54 / 300 = **18.0%** | 250 / 1,000 = **25.0%** | So: * A: **230 / 1,000 = 23.0%** * B: **250 / 1,000 = 25.0%** * Absolute difference: **+2.0 percentage points** * Relative conversion-rate change: **(25% − 23%) / 23% = 8.70%** Therefore, the announcement's **“B improves signup conversion by 8.7%”** is mathematically equivalent to the observed aggregate rate difference, **if “improves” is intended descriptively as “had an 8.7% higher observed conversion rate.”** But that wording can misleadingly imply causation. ### 2. The important complication: device mix The designs have very different device distributions: * A: **30% desktop / 70% mobile** * B: **70% desktop / 30% mobile** And desktop converts substantially better than mobile **for both designs**: * A: 30% desktop vs. 20% mobile * B: 28% desktop vs. 18% mobile Within **each device group, B is actually lower**: * Desktop: B **28%** vs. A **30%** → **−2 pp** * Mobile: B **18%** vs. A **20%** → **−2 pp** So the aggregate result reverses because B received a much larger share of desktop users. This is a classic composition/mix issue. In this particular dataset, if the two designs had the **same device mix**, B would not have the observed aggregate advantage. For example, applying A's 30% desktop / 70% mobile mix to B's device-specific rates: **0.30 × 28% + 0.70 × 18% = 21.0%** compared with A's **23.0%**. Conversely, applying B's 70% desktop / 30% mobile mix to A: **0.70 × 30% + 0.30 × 20% = 27.0%** compared with B's **25.0%**. That strongly illustrates why the overall 25% vs. 23% comparison should not be interpreted as a design effect without understanding assignment. ### 3. What I would do next I would **not announce an all-user rollout yet**. First establish how users were assigned. Specifically, check: 1. **Assignment mechanism:** Were users randomly assigned to A/B? 2. **Allocation over time:** Did the A/B split change during the 14 days? 3. **Targeting rules:** Were particular devices, countries, traffic sources, account types, or other audiences preferentially exposed to one design? 4. **Exposure counts over time:** Confirm whether the device imbalance was present throughout the period or arose from a change in allocation. 5. **Statistical uncertainty:** Calculate confidence intervals or an appropriate test for the comparison, ideally both overall and within device strata. 6. **Pre-specified primary metric:** Confirm that seven-day completed signup was the intended primary outcome and that the analysis wasn't selected after seeing the results. The key question is whether the assignment process makes the designs comparable. If assignment was properly randomized and balanced/appropriately analyzed, the evidence can support a causal comparison. If assignment was targeted or otherwise non-random, the observed difference could reflect who received each design rather than the designs themselves. ### Decision note Here is a version you can take to the team: **Signup-page A/B results — decision note** **Summary** The export shows an observed seven-day signup conversion rate of **25.0% for Design B (250/1,000)** versus **23.0% for Design A (230/1,000)**. This is a **2.0 percentage-point absolute difference**, or an **8.7% relative difference**. However, these aggregate results do not currently establish that Design B caused the improvement or support an immediate rollout to all users. **Why** The device mix differs substantially between the designs: * Design A: 300 desktop and 700 mobile users * Design B: 700 desktop and 300 mobile users Within both device groups, B has a lower observed conversion rate than A: * Desktop: A 30.0% vs. B 28.0% * Mobile: A 20.0% vs. B 18.0% Thus, the aggregate advantage for B is driven by its much higher proportion of desktop users, whose conversion rate is higher under both designs. The aggregate 25.0% vs. 23.0% result should therefore not be treated as a clean estimate of the design effect without examining assignment and device composition. **What we should verify before rollout** 1. How users were assigned to A or B, including whether assignment was randomized. 2. Whether the allocation between designs changed during the 14-day period. 3. Whether targeting rules caused different user populations to receive the two designs. 4. Whether the device mix and conversion differences were consistent over time. 5. Statistical uncertainty around the observed differences, including device-level comparisons. **Current decision** Do not characterize the result as evidence that B “improves” conversion by 8.7% in a causal sense yet, and do not use these aggregate numbers alone as the basis for an all-user rollout. The verified descriptive result is: **B had a 25.0% observed seven-day signup conversion rate versus 23.0% for A during this 14-day export, a relative difference of 8.7%.** A rollout decision should follow the assignment/targeting review and an appropriately stratified statistical analysis. The most important takeaway for the team is that **the 8.7% calculation is correct as an aggregate relative difference, but the proposed causal interpretation and rollout do not follow from these numbers alone.**
Gemini · 2026-09-24T00:12:06.020064+00:00
I'm a product manager reviewing two signup-page designs. Our team wants to ship design B to everyone because the dashboard says its signup conversion rate is higher. Can you analyze the results, tell me what you would do next, and draft the decision note I can take to the team? This is a fictional exercise. The export covers the same 14-day enrollment period for both designs, and every included user has a complete seven-day follow-up. Each row counts distinct eligible users who first saw that design, with completed signups counted once within seven days of that exposure. A user appears in only one design and one device group; device is recorded at first exposure. Desktop and mobile are the only groups. There are no missing counts in this export. I have not yet checked how users were assigned to designs, whether allocation changed during the period, or how targeting was configured. Design A, desktop: 300 exposed users, 90 completed signups. Design A, mobile: 700 exposed users, 140 completed signups. Design B, desktop: 700 exposed users, 196 completed signups. Design B, mobile: 300 exposed users, 54 completed signups. The proposed announcement says: "B improves signup conversion by 8.7%, so we should roll it out to all users." I need to know whether that claim and rollout decision follow from these numbers. Please show the calculations that matter and explain what evidence we still need. The proposed claim and rollout decision are **incorrect** and based on a classic statistical illusion: **Simpson's Paradox**. While Design B appears to win overall, Design A is actually equal or superior on both desktop and mobile when analyzed separately. --- ### Step 1: The Calculations That Matter #### 1. Segment-by-Segment Conversion Rates * **Desktop:** * Design A: $90 / 300 =$ **30.0%** * Design B: $196 / 700 =$ **28.0%** * *Verdict:* **Design A wins** by 2.0 percentage points. * **Mobile:** * Design A: $140 / 700 =$ **20.0%** * Design B: $54 / 300 =$ **18.0%** * *Verdict:* **Design A wins** by 2.0 percentage points. #### 2. Aggregate Conversion Rates * **Design A Overall:** $(90 + 140) / (300 + 700) = 230 / 1,000 =$ **23.0%** * **Design B Overall:** $(196 + 54) / (700 + 300) = 250 / 1,000 =$ **25.0%** * *Difference:* $25.0\% - 23.0\% = +2.0$ percentage points ($+8.7\%$ relative improvement). #### Why is this happening? The aggregate metric is distorted because the **traffic mix (device allocation) is heavily unbalanced**: * Desktop users convert at a much higher baseline rate (~28–30%) than mobile users (~18–20%). * **Design B received 70% desktop traffic**, while **Design A received only 30% desktop traffic**. * Design B looks better overall simply because it was shown to significantly more high-converting desktop users, not because the design itself performs better. --- ### Step 2: What Evidence We Still Need Before making any rollout decision, you need to investigate **why the allocation was skewed**: 1. **Traffic Split Verification:** Was this an unstratified 50/50 split that accidentally suffered from a bad random assignment, or was targeting misconfigured (e.g., routing desktop users to Design B)? 2. **Allocation Over Time:** Did the split change mid-experiment? (e.g., ramping up Design B on desktop midway through). 3. **Statistical Significance:** With $N = 1,000$ per variant, running a chi-squared test or two-proportion z-test on each segment will reveal whether Design A’s 2 percentage point lead on mobile and desktop is statistically significant or just noise. 4. **Behavioral Metrics Post-Signup:** Since conversion rates within device types favor Design A, check down-funnel retention or activation to ensure user quality isn't compromised. --- ### Step 3: What To Do Next 1. **Do NOT ship Design B to everyone.** Shipping Design B globally would likely decrease your overall conversion rate due to its lower segment performance. 2. **Pause or fix the experiment setup** to ensure an even 50/50 randomized split across both device types (stratified sampling). 3. **Re-run or continue the test** with balanced allocation until statistical significance is reached within each device segment. --- ### Step 4: Draft Decision Note for the Team ```text SUBJECT: Hold on Design B Rollout – Re-evaluating A/B Test Results Team, I am recommending that we pause the full rollout of Design B. While the top-line dashboard shows Design B with a 25.0% overall conversion rate vs. Design A's 23.0% (+8.7% relative lift), a segment breakdown reveals that this is caused by Simpson's Paradox due to an uneven traffic split across devices. Key Findings: 1. Desktop Conversion: Design A (30.0%) beats Design B (28.0%). 2. Mobile Conversion: Design A (20.0%) beats Design B (18.0%). 3. Traffic Imbalance: Design B was exposed to 70% desktop traffic (which converts at a higher baseline rate), while Design A was exposed to 70% mobile traffic. Because Design A actually outperforms Design B on both desktop and mobile individually, rolling out Design B to 100% of users would likely cause our total signups to drop. Next Steps: - Engineering/Data Science to review experiment configuration to understand why traffic allocation was skewed (70/30 instead of 50/50 per device). - Calculate statistical significance on the per-device segment results. - If necessary, run a clean follow-up test with proper stratified 50/50 randomization before making a final decision. Let's discuss during our sync today. ```
Claude · 2026-09-24T00:12:06.068169+00:00
I'm a product manager reviewing two signup-page designs. Our team wants to ship design B to everyone because the dashboard says its signup conversion rate is higher. Can you analyze the results, tell me what you would do next, and draft the decision note I can take to the team? This is a fictional exercise. The export covers the same 14-day enrollment period for both designs, and every included user has a complete seven-day follow-up. Each row counts distinct eligible users who first saw that design, with completed signups counted once within seven days of that exposure. A user appears in only one design and one device group; device is recorded at first exposure. Desktop and mobile are the only groups. There are no missing counts in this export. I have not yet checked how users were assigned to designs, whether allocation changed during the period, or how targeting was configured. Design A, desktop: 300 exposed users, 90 completed signups. Design A, mobile: 700 exposed users, 140 completed signups. Design B, desktop: 700 exposed users, 196 completed signups. Design B, mobile: 300 exposed users, 54 completed signups. The proposed announcement says: "B improves signup conversion by 8.7%, so we should roll it out to all users." I need to know whether that claim and rollout decision follow from these numbers. Please show the calculations that matter and explain what evidence we still need. # Bottom line **The claim and the rollout decision do not follow from these numbers.** B's higher overall rate comes from a difference in device mix between the arms, not from B performing better. Within each device group, B converts 2.0 points lower than A. The data is also too small to show B is worse. ## 1. Overall result (what the dashboard shows) | Design | Exposed | Signups | Rate | |---|---|---|---| | A | 1,000 | 230 | 23.0% | | B | 1,000 | 250 | 25.0% | - Absolute difference: +2.0 pp. Relative lift: 25.0 / 23.0 − 1 = **8.7%**, which is where the announcement's figure comes from. - The claim calls this "conversion improves by 8.7%", but it is a relative lift, and the absolute change is 2.0 pp. - Even taken at face value, this is not statistically distinguishable from zero. Pooled z ≈ 1.05, p ≈ 0.30, 95% CI for the difference ≈ −1.7 to +5.7 pp. ## 2. By device (the reversal) | Device | A rate | B rate | B − A | 95% CI for B − A | |---|---|---|---|---| | Desktop | 90/300 = 30.0% | 196/700 = 28.0% | −2.0 pp | −8.2 to +4.2 pp | | Mobile | 140/700 = 20.0% | 54/300 = 18.0% | −2.0 pp | −7.3 to +3.3 pp | B is lower in both groups yet higher overall. This is Simpson's paradox, driven by traffic mix: | Design | Desktop share | Mobile share | |---|---|---| | A | 30% | 70% | | B | 70% | 30% | Desktop converts better for both designs, and B's traffic is 70% desktop while A's is 70% mobile. B's overall number is lifted by who saw it, not by the design. ## 3. Same-mix comparison Weighting both designs to the same device mix removes the composition effect: | Standard mix | A | B | B − A | |---|---|---|---| | A's mix (30% desktop) | 23.0% | 21.0% | −2.0 pp | | B's mix (70% desktop) | 27.0% | 25.0% | −2.0 pp | | 50/50 | 25.0% | 23.0% | −2.0 pp | The device-adjusted estimate is **−2.0 pp** (B lower), with a 95% CI of roughly −6.1 to +2.1 pp (z ≈ −1.0). So the data does not show B is worse either. It is inconclusive. ## 4. A second red flag: the allocation itself Each design has 1,000 users, but the device split is 300/700 versus 700/300. If users were randomly assigned to A or B regardless of device, the splits would be near-identical. A chi-square test on device × design gives χ² = 320, which is far beyond chance. **Assignment was not independent of device**, so this does not look like a clean randomized test. Possible causes are device-based targeting, a change in allocation partway through the period, or a rollout by platform. This matters because the device-adjusted comparison in section 3 is only trustworthy if assignment within each device group was random. That is unverified. ## 5. Evidence still needed 1. **Assignment mechanism:** who or what decided A vs B, and whether it was random within each device group. 2. **Allocation history:** daily exposure counts by design and device, to see whether allocation or targeting changed mid-period. If it did, time is another confounder. 3. **Targeting rules:** any device, geography, traffic source, or new/returning filters that differed by design. 4. **Other mix differences:** acquisition channel, country, browser, and campaign timing by design, since device may not be the only imbalance. 5. **Sample ratio check:** the intended split versus the observed split, overall and by device. 6. **Guardrails:** signup quality, activation, retention, error rates, and page speed. A signup lift that doesn't carry through to downstream metrics isn't a win. 7. **Power:** at a ~23% baseline, detecting a 2 pp difference needs roughly 7,000 users per arm (80% power, α = 0.05). This export has 1,000 per arm, and per-device analysis needs more. ## 6. What I would do next 1. Do not announce or roll out B based on this dashboard. 2. Get the assignment and allocation answers above, since they decide whether this data can be salvaged. 3. Rerun as a properly randomized test, stratified by device, with a fixed sample size and a pre-declared primary metric (device-adjusted signup rate) and guardrails. 4. If the two designs need to be compared sooner, use the device-adjusted comparison as the primary readout, not the pooled dashboard number. ## Draft decision note **Title:** Signup page B vs A: rollout decision **Status:** Not ready to roll out. Recommend a controlled rerun. **Rationale:** - The dashboard shows B at 25.0% vs A at 23.0% (+2.0 pp, +8.7% relative). This gap is not statistically significant (p ≈ 0.30). - The overall gap is a device-mix artifact. B's traffic was 70% desktop and A's was 70% mobile, and desktop converts better for both designs. - Within each device, B is lower: desktop 28.0% vs 30.0%, mobile 18.0% vs 20.0%. Adjusted for device mix, B − A is about −2.0 pp (95% CI ≈ −6.1 to +2.1 pp), which is inconclusive. - The device split between designs is far from what random assignment would produce, so assignment and targeting are unverified. **Cost:** [Engineering and design cost of rollout: to be filled in]. Cost of delay: one more test cycle. Cost of a wrong rollout: possible lower conversion and an announcement we would have to retract. **Expected effect:** Unknown. The data supports neither a B improvement nor a B decline; the plausible range for B's effect spans roughly a 6 pp decrease to a 2 pp increase. **Response:** 1. Hold the "8.7% improvement" announcement and the full rollout. 2. Confirm the assignment mechanism, allocation history, and targeting rules. 3. Rerun with device-stratified randomization and a pre-set sample size (about 7,000+ per arm for a 2 pp effect), tracking signup quality and retention as guardrails. 4. Decide on the device-adjusted result. **Assumptions:** I treated the export as accurate and complete as you described it. The confidence intervals use normal approximations, and the device-adjusted estimate assumes assignment was random within each device group, which is exactly what we still need to verify.