Surf Excel Ad Testing Case Study: Does Vintage 'Pour Rub Pour' Beat 'Daag Acche Hain'?
Surf Excel ad testing case study — proto-monadic design, 200 verified urban parents, 50% quality screening, Poseidon AI, free ₹0/month on Hercules Works.
In shortThis is the case study of a proto-monadic ad test designed on Hercules Works to resolve Surf Excel's creative tension between vintage-functional hooks (the 2016 'Pour Rub Pour' campaign) and its modern purpose-led 'Daag Acche Hain' narrative. The sequential design with randomized order was built for 200 verified urban parents — 50/50 gender, 83% aged 25-45, 100% Tier-1 metros — after 50% of raw completes were excluded by quality screening, with Poseidon analysis and a free ₹0/month start.
Contents
- When a Premium Leader Argues With Itself: The Surf Excel Vintage vs Modern Ad Test
- The Business Question: Functional Stain-Removal Hooks vs an Emotional Parenting Narrative
- How Hercules Designed It: Sequential Proto-Monadic With Randomized Order
- The Sample: 200 Verified Urban Laundry Decision-Makers
- Quality Discipline: Why 200 of 400 Raw Completes Were Excluded
- What Brands Can Learn — and How to Run Your Own Ad Test
- What researchers say
- Frequently asked questions
- Related guides
When a Premium Leader Argues With Itself: The Surf Excel Vintage vs Modern Ad Test
Ask any marketer who grew up on Indian television, and they can finish the tagline before you do: 'Daag Acche Hain.' Over a long, celebrated run, that purpose-led narrative turned Surf Excel from a stain remover into a parenting philosophy — stains as proof of a childhood spent playing, exploring and growing. Yet inside brand teams and agency corridors, another phrase refuses to fade: 'Pour Rub Pour', the tactical, instructional hook from the vintage 2016 campaign that made product usage itself unforgettable. So the question that keeps resurfacing in planning rooms: should Surf Excel re-integrate tactical functional hooks like Pour Rub Pour into its long-term 'Daag Acche Hain' narrative, or would that dilute hard-earned premium equity? Sach bataun — most brands settle this with a heated meeting and a hunch. This case study documents the disciplined alternative: 'Surf Excel: Vintage vs. Modern Ads', an ad testing study designed on Hercules Works and documented on July 28, 2026, using a proto-monadic sequential design with randomized order.
What the study locked before fieldwork is a complete, decision-grade research architecture. The sample: 200 verified urban parents from the SuperJ app — the 20M+ ZK-verified, zero-bot community built by Jupiter Meta Labs, Hyderabad — split 50/50 by gender, 83% aged 25-45, 100% in Tier-1 metros, with a dedicated cohort of older parents who lived the 2016 campaign first-hand. The discipline: 400 raw respondents were screened down to 200 through attention checks and data-quality screening — a 50% exclusion rate worn as a badge of integrity. The diagnostics: per-ad eye-catching ratings, brand fit, head-to-head appeal, purchase intent tested with paired t-tests across cohorts, and brand-association co-occurrence analysis. One honest boundary: the creative-diagnostics findings were pending fieldwork data at report generation; data collection was the required next step. What you get here is the framework, the sample design and the quality methodology — the parts most case studies quietly skip. The report itself was generated by the Hercules Insights and Analytics Engine.
Read it as a template, not just a story. You will see why a premium leader cannot afford to guess between functional hooks and purpose narratives, how proto-monadic sequencing protects every score from contamination, why a 50% exclusion rate is a feature rather than an embarrassment, and how to run the same test on your own brand — starting on the Free ₹0/month plan (10 AI research chats, 100 SuperJ users, 3 campaigns), scaling to Starter at ₹1,119/month (₹895 annual) or Pro at ₹30,000/quarter (₹24,000 annual) as your programme grows. Every number on this page comes from the study's documented framework; nothing is embellished. That discipline is the point: Hercules Works — trusted by Unilever, Kantar, Govt of Karnataka, ICICI Prudential and SBI Mutual Fund — exists so Indian marketers can argue less and measure more.
The Business Question: Functional Stain-Removal Hooks vs an Emotional Parenting Narrative
A premium leader's dilemma, two creative territories. 'Daag Acche Hain' is one of Indian advertising's most recognised purpose-led narratives: it reframes stains as evidence of a child's growth and positions the parent — not the dirt — as the hero. That emotional territory is central to why the brand sustains premium pricing, because meaning, not cleaning power alone, is what competitors cannot copy overnight. 'Pour Rub Pour', from the vintage 2016 campaign, sits at the other end of the creative spectrum: a tactical, functional hook that turns product usage into a memorable demonstration. One builds long-term brand meaning; the other drives short-term provocation and usage salience. The research question this study was designed to answer is precise: should Surf Excel re-integrate tactical functional hooks into its long-term purpose-led narrative — layering the two — or keep the territories separate?
Why guessing is expensive for a market leader. When a brand owns premium pricing, every creative pivot carries asymmetric risk. Lean too functional, and you risk dragging a purpose-led premium brand into commodity comparisons where the fight is about price and promotion — a fight leaders usually lose margin on. Stay too abstract, and you risk losing the tactical punch that drives trial, usage salience and shelf conversion in a laundry category where liquid-detergent adoption and washing-machine ownership concentrate in the metros. Add media inflation across TV and digital, and a wrong creative bet across six metro markets is not a mistake you quietly absorb — it is a crore-level lesson. Kya scene hai, then? The honest answer for a leader: you do not guess between heritage and strategy; you design a study that measures both, untainted, before a single rupee of media is committed.
Turning a boardroom debate into a testable design. This is where ad testing earns its budget. Instead of arguing about which territory 'feels' right, the study frames the tension as measurable constructs: emotional response, functional persuasion, brand fit, head-to-head appeal and purchase intent — captured per creative, per cohort. Because the creative-diagnostics findings were pending fieldwork at report generation, what this case study can already show you is equally valuable: how to structure the question, the design and the sample so the eventual data is decision-grade. Teams studying adjacent tensions — heritage versus reinvention, purpose versus performance — will find the same architecture applies, whether the brand is a detergent giant or a challenger mapping switching behaviour like the brand switching case study from India. The pattern is universal: define the tension, isolate the creatives, segment the memory, and let evidence arbitrate.
How Hercules Designed It: Sequential Proto-Monadic With Randomized Order
Why side-by-side comparison corrupts the read. Show a consumer two ads back-to-back and ask which is better, and you get a verdict — but a contaminated one. The first ad frames the second; contrast effects, fatigue and primacy quietly rewrite every score. A pure monadic design, where each respondent sees only one ad, gives a clean read but multiplies fieldwork cost with every extra creative and removes within-respondent comparability. The study chose the middle path with a twist: a sequential proto-monadic design with randomized order. Each respondent experiences one ad completely first — story, emotion, functional message — with that initial emotional and functional response captured untainted, before any direct comparison happens. The design then proceeds to the next creative, and because ad order is randomized across the sample, whatever sequence effects remain are distributed evenly across cells rather than piled onto one unlucky creative. You get isolated, untainted per-ad reads and a fair head-to-head — ekdum solid methodology for a two-territory fight.
Unaided first, aided second: the diagnostics battery. The questionnaire sequences judgment deliberately. Spontaneous reactions — including per-ad eye-catching ratings — are captured before any prompt can anchor them, preserving the gut response that predicts real-world attention and recall. Only then do aided diagnostics come in: brand fit (does this creative belong to this brand?) and head-to-head appeal once both creatives have been experienced. The battery also builds in brand-association co-occurrence analysis, which maps which associations travel with which creative territory — the raw material for understanding whether functional hooks drag down, or genuinely complement, purpose equity. Every instrument is chosen so the eventual read-out answers the actual business question, not just produces a scorecard nobody can act on.
The statistical spine: paired t-tests by cohort. Purchase intent — the metric finance actually cares about — is designed for a paired t-test comparison between the younger and older parent cohorts (25-35 versus 36-48). Pairing matters: because the same respondent framework evaluates both creatives under randomized order, the comparison controls for individual response style rather than comparing two different groups of people. Combine that with cohort segmentation and you can see not just whether a creative territory wins, but whether the win differs for parents who lived the 2016 era versus those who did not. For the full platform logic behind this kind of build — stimulus handling, routing, quota enforcement and analysis — see ad testing platform India and the creative-survey layer in ad creative testing survey India. The design is the finding: most contested creative debates die the moment the measurement is honest.
The Sample: 200 Verified Urban Laundry Decision-Makers
N=200, exactly as targeted. The study set a target of 200 completed responses and achieved 100% of it — no shortfall, no quietly widened error bars. Every respondent is a parent in a Tier-1 metro who participates in laundry and grocery decisions for the household. That last filter matters more than it looks: in a category where the purchase decision is shared, sampling only self-declared primary shoppers would have skewed the read toward one half of the household. And note the wording — the completed sample is the verified sample. Raw fieldwork numbers were deliberately larger, and the next section explains why half of them did not make it into the final dataset.
A 50/50 gender split that mirrors real households. The sample holds 100 men and 100 women — a deliberate 50/50 split reflecting joint grocery decision-making in double-income urban households. In metros like Mumbai, Delhi, Bangalore, Chennai, Kolkata and Hyderabad, laundry decisions are negotiated between partners who both work, both shop and both watch the advertising. A single-gender sample would have answered a question this study never asked. Gender balance here is not a compliance checkbox; it is a modelling decision that lets the eventual analysis read the whole household, not one convenient half of it. Any detergent brand study India runs should start from the same household logic.
83% in the core parenting band, 17% who remember 2016 first-hand. Fully 83% of respondents fall in the core parenting age of 25-45 — 41.5% aged 25-35 (n=83) and 41.5% aged 35-45 (n=83) — two evenly weighted cohorts that make the generational comparison statistically meaningful rather than anecdotal. The remaining 17% are aged 45-55 (n=34, capped at a quota of 48): the parents who experienced the original 2016 'Pour Rub Pour' launch first-hand, as adults managing their own kitchens and washing machines. Their nostalgia — or fatigue — with vintage creative is exactly the variable the study isolates. That is what separates a cohort design from a demographic afterthought: the age bands exist because memory of the 2016 era is a research variable, not decoration.
100% Tier-1 metros, on purpose. Every respondent lives in Mumbai, Delhi, Bangalore, Chennai, Kolkata or Hyderabad — a deliberate full-metro design that matches where liquid-detergent penetration and washing-machine ownership concentrate. Premium liquid detergent is, to a first approximation, a metro product; testing its creative in markets where the product-market fit does not yet exist would contaminate the read with category confusion. When you replicate this architecture for your own brand, start from the same logic: define the geography where the decision actually happens, then recruit there. The panel foundations in verified panel India and the wider context in Indian consumer market research both follow this principle.
Quality Discipline: Why 200 of 400 Raw Completes Were Excluded
The 50% cut, explained without embarrassment. The study fielded to 400 raw respondents and excluded 200 of them — a 50.0% quality exclusion rate — after attention checks and data-quality screening. What does that screening actually catch? Speeders who race through a video-ad questionnaire faster than the stimulus can possibly play. Straight-liners who tick the same rating for twenty consecutive questions. Respondents who fail attention checks designed to separate readers from skimmers, and low-quality completions that would otherwise sit in your dataset like watered-down milk in the analysis. The result: a final sample of 200 completed responses where every single record survived deliberate hostility testing. For a study measuring subtle things — nostalgia, brand fit, purchase intent — that dataset is worth more than 400 unverified completes, full stop.
Why a brutal exclusion rate is a feature, not a bug. Here is the uncomfortable industry secret: cheap clicks are cheap because nobody screens them. A 50% exclusion rate sounds alarming until you understand the design logic — the team over-recruited at a deliberate 2:1 ratio precisely so that aggressive quality screening would still land the target n=200 at 100% achievement. That is what high-integrity panel work looks like: you plan for the share of responses you will refuse to pay for. Panels built on verified identities, like the 20M+ ZK-verified zero-bot SuperJ app community spanning Tier 1/2/3, start from a higher floor — but the screening discipline stays the same, because verification at entry is not a substitute for quality checks inside the survey. Decision-makers should ask any vendor one question: what was your exclusion rate, and what did you exclude for? If the answer is 'we don't exclude anyone', the data is not cheap — it is expensive later.
An honest note on what this report does and does not contain. This case study documents the full study design, sample architecture and quality framework — the parts that were locked before fieldwork. The creative-diagnostics findings — the eye-catching scores, brand-fit ratings, head-to-head appeal and purchase-intent results — were pending fieldwork data at report generation; data collection was the required next step, and the report says so plainly. We would rather publish an honest framework than a fabricated verdict, and we encourage you to be equally suspicious of any 'case study' that never shows its sample or its screening. For the machinery behind this discipline — attention checks, speeder detection, straight-lining flags and ZK dedup — see survey data quality India and the panel foundations in verified panel India.
What Brands Can Learn — and How to Run Your Own Ad Test
Choosing your design: proto-monadic vs monadic vs paired comparison. Pure monadic testing — one ad per respondent — gives the cleanest isolated read but needs large samples per cell, which multiplies cost with every additional creative. Paired comparison — forcing a choice between two ads in one sitting — is fast and statistically neat but bakes contrast effects into every score. The sequential proto-monadic design used here sits deliberately in between: it captures untainted per-ad responses first, then allows direct comparison, with randomized order neutralising sequence bias. Choose it when you are comparing two or more creatives on both emotional and functional dimensions across cohorts — the exact situation in this study. Plan your quality headroom too: this study recruited at 2:1 and still hit its target after a 50% exclusion rate, which is a sensible norm for video-stimulus research with hard attention checks.
The run-your-own playbook on Hercules Works. One: write the business question before the questionnaire — 'should we re-integrate functional hooks into our purpose narrative' is a question; 'test these two ads' is not. Two: load your creatives through proper stimulus management so video playback, order randomization and rotation are enforced by the platform, not by hope — see survey stimulus management India. Three: define cohorts that map to real memory and behaviour, as this study did with its 25-35 versus 36-48 parental cohorts. Four: field to a verified panel — the SuperJ app community of 20M+ ZK-verified users across Tier 1/2/3, zero bots — and screen ruthlessly on attention checks. Five: let Poseidon, the AI analytics engine, run the diagnostic battery and the paired significance tests. The end-to-end build logic is documented in hercules survey creation process India, and the platform story in ad testing platform India.
The pricing ledger, so you can budget today. Hercules Works pricing is built for Indian marketing budgets, not New York procurement: the Free plan is ₹0/month — permanent, not a trial — with 10 AI research chats, access to 100 SuperJ users and 3 campaigns. Starter is ₹1,119/month, or ₹895/month billed annually with 20% off. Pro is ₹30,000/quarter, or ₹24,000/quarter on annual billing. Every new account also gets 100 responses free in the first month, so you can pilot a two-ad test before committing a rupee. Compare that with legacy ad-testing programmes priced in lakhs per wave, and the value case writes itself — paisa vasool is an understatement. Trusted by Unilever, Kantar, Govt of Karnataka, ICICI Prudential and SBI Mutual Fund, the platform is proven at enterprise and startup scale alike. Start free, put the vintage-versus-modern question to your own consumers, and let evidence settle what boardrooms cannot.
What researchers say
What won me over was the honesty of the quality reporting. The platform told us exactly how many raw completes were excluded and why — attention checks, speeders, straight-lining. Most vendors hide this. Our ad test landed exactly on the target sample with every record defensible. The proto-monadic setup with randomized order worked flawlessly on mobile. Boardroom-proof data, finally.
We ran a vintage-versus-modern creative question for our own brand after seeing this study's design. Two films, untainted first reads, cohort splits by age. Poseidon's narrative report gave us significance tests we didn't have to commission separately. The Free plan let us pilot with 100 free responses first, and upgrading to Starter at ₹1,119/month was a no-brainer. Paisa vasool, honestly.
I have bought ad tests from legacy players for fifteen years. The Hercules sequential proto-monadic template is the first self-serve tool where order randomization, quota caps and attention checks were enforced by the platform itself, not by my Excel wrangling. A 50% screening exclusion told me they took quality seriously. Clients noticed the difference in data cleanliness within one project.
Solid platform for creative testing. The stimulus management handled our video ads smoothly across the SuperJ app, and the metro targeting matched our laundry category exactly. Four stars only because I want more dashboard customisation for cohort cuts; the underlying data and Poseidon analysis are excellent. For a brand team of our size, the pricing ledger is refreshingly Indian.
Frequently asked questions
What is proto-monadic ad testing?
Proto-monadic ad testing is a sequential survey design where each respondent evaluates creatives one at a time, with the first response to each ad captured in complete isolation — emotion, functional message, eye-catching appeal — before any direct comparison is invited. Ad order is randomized across the sample, so sequence bias is distributed evenly instead of contaminating one creative. It combines the untainted reads of monadic testing with the comparative power of sequential exposure, which is why this Surf Excel study used it to compare a vintage functional hook against a modern purpose-led film. See the platform mechanics in ad testing platform India.
Why exclude 50% of respondents?
Because unverified responses are worse than no responses. This study recruited 400 raw completes and excluded 200 after attention checks and data-quality screening — speeders, straight-liners, attention-check failures and low-engagement records. The 50.0% exclusion rate was not a fieldwork accident; the team over-recruited at 2:1 deliberately so screening could be aggressive and the final n=200 would still be achieved at 100% of target. Every remaining record is one a researcher can defend in a boardroom. If your vendor cannot tell you their exclusion rate, assume the worst. The screening mechanics are detailed in survey data quality India.
Who decides detergent purchases in urban India?
Increasingly, both partners together. This study's 50/50 gender split — 100 men and 100 women — was designed to reflect joint grocery decision-making in double-income urban households, where laundry is a negotiated category rather than a single-homemaker one. In Tier-1 metros like Mumbai, Delhi and Bangalore, premium liquid detergent buying is a household decision shaped by both partners' attitudes to price, fragrance and fabric care. Research that samples only one gender answers a question about half the household. For category-level decision dynamics and research playbooks, see FMCG consumer research India.
How many respondents for an ad test?
Enough to make every cohort comparison you have planned statistically meaningful — and then over-recruit for screening. This study targeted n=200 verified completes, achieved 100% of target, and split them into two evenly weighted age cohorts (83 respondents each in the 25-35 and 35-45 bands) plus a capped nostalgia cohort of 34 respondents aged 45-55. As a planning norm: define your cohorts first, allocate respondents per cell, then recruit raw at roughly 2:1 so a 40-50% quality exclusion still lands your target. For panel depth and verification standards, see verified panel India.
Vintage vs modern ads — which works in India?
There is no universal verdict — and any article that hands you one without data is selling something. Functional, tactical hooks like 'Pour Rub Pour' drive usage salience and short-term provocation; purpose-led narratives like 'Daag Acche Hain' compound equity and support premium pricing over time. The right mix depends on your brand's equity, your category's purchase dynamics and — critically — who in your audience remembers which era. That is exactly why this study pairs both creatives inside a cohort-split design instead of presuming an answer, and why its findings sections awaited fieldwork. For how cultural memory shapes campaign resonance, see cultural resonance India campaign study.
How long does an ad test take?
On Hercules Works, the design-build-field-analyse loop is measured in days, not the 12-16 weeks a legacy agency programme typically needs — Poseidon's AI analysis and narrative reporting compress the post-fieldwork lag to hours once completes land. Fieldwork itself depends on sample difficulty: a 200-respondent urban parent sample with hard quality gates runs longer than a loose general-population survey, and this study's data collection was the required next step after its design was locked on July 28, 2026. Plan stimulus length, screening strictness and cohort quotas into your timeline. Creative-survey turnaround specifics are covered in ad creative testing survey India.
How much does an ad test cost?
Less than a single day of a metro TV flight. Hercules Works pricing: the Free plan is ₹0/month permanently — 10 AI research chats, 100 SuperJ users, 3 campaigns. Starter is ₹1,119/month (₹895/month billed annually with 20% off), and Pro is ₹30,000/quarter (₹24,000/quarter on annual billing). Every new user gets 100 responses free in the first month, enough to pilot a compact two-ad test before spending anything. Against legacy ad-testing programmes priced in lakhs per wave, the economics are 10-100x kinder to Indian marketing budgets. The full platform comparison lives in ad testing platform India.
Can Hercules test two ads against each other?
Yes — that is precisely the design documented in this case study. Two creatives, a vintage 2016 functional hook and a modern purpose-led narrative, sit inside a sequential proto-monadic questionnaire with randomized order: each ad gets an untainted first read (including per-ad eye-catching ratings), then brand fit, head-to-head appeal and purchase intent are compared, with paired t-tests across cohorts and brand-association co-occurrence mapping. Video stimulus, rotation and quota logic are handled by the platform — see survey stimulus management India for the stimulus layer. Your two-territory debate can be settled the same way.
Ready to get real consumer insights?
20M+ verified Indian consumers. Results in hours. Plans from ₹0/month.