Confessions of a Data Point
[marketing
]
The first iPod I ever owned, a silver iPod Classic that I loved beyond all reason, was paid for by consumer research panels in Mumbai. I was eighteen or nineteen and in college, and every few weeks an agent would call to say there was a study happening at a nice hotel, that it paid somewhere between two and five thousand rupees, and that all I had to do was show up and talk about products for a couple of hours. There would be good food, a mix of new and familiar faces, genuinely interesting questions, and cash at the end of it. For a college student this was about as close to a perfect gig as the world offered, and I would love to tell you that I thought carefully about the integrity of the research I was participating in, but mostly I thought about the iPod. Totally worth it imo.
The problem, which I understood only dimly at the time, was that I was almost never the person those studies were actually looking for. The screener might want women in their thirties who had recently switched detergent brands, or young professionals earning a certain amount, and I was a student earning nothing at all, and yet the agent would pitch me in anyway. What he did check, carefully and before anything else, was whether you could hold a conversation. There was always a sort of vibe check first: could you communicate well, could you sound charming, would you be pleasant to have in a room for two hours. If you passed that, the demographic requirements had a way of becoming negotiable.
Once you were in the room the negotiation continued, sometimes quite far past what the study designers would have found reassuring. I remember being pulled into a study for a new mouthwash where they very carefully would not tell us the brand name, presumably so that our reactions would not be biased, and where we were expected to actually swish and rinse with the product they handed us. I could not bring myself to do it. I had no idea what was in the bottle, whether it had been through the kind of testing that would let a random college student put it in her mouth without concern, and the money was not enough to override that particular caution, so I simply made up my answers. I described a taste, I described a level of freshness, I rated the burn, and I handed the form back. Whoever wrote the report on that mouthwash carries, somewhere in his findings, my invented opinion of a product I never tasted.
A friend of mine had a stranger version of the same thing happen. He was pitched into a panel for cigarettes even though he had never smoked one in his life, which the agent either did not know or did not particularly care about, and part of the protocol required participants to actually try the product. Everyone was ushered outside for the smoking portion. My friend, faced with the choice between refusing the study, admitting he had lied to qualify, or lighting a cigarette for the first time in his adult life, picked a fourth option: he lit it, stubbed it out almost immediately, waited a polite amount of time with the smokers, and came back inside to report on the experience. His feedback, whatever it was, is presumably still sitting in a slide somewhere on how habitual smokers responded to that particular blend.
I do not know exactly where any of these answers went. This was consumer research, so presumably the data flowed to brand managers somewhere trying to figure out how to launch a new product or how to shift how people felt about an existing one. Who knows. What I do know is that the sample they believed they were studying and the room they actually got were two different things, and that the difference was invisible to everyone downstream. The transcripts would have read beautifully, because we had been selected, above everything else, for our ability to talk.
That last part is the piece I keep coming back to, now that I have spent years on the other side of the one-way glass. Nobody analyzing that data could have caught us. The quotas were filled, the demographics were recorded as whatever the recruiter had written down, and the discussion was lively and articulate, because articulateness was the one criterion that had been genuinely enforced. A careful analyst could have run every consistency check available and found nothing wrong, because the corruption did not live in the data. It happened before the data existed, at the moment of collection, and it left no statistical trace whatsoever.
Whose incentives were those, exactly?
It would be comfortable to file all of this under one dishonest recruiter in one city, but I think that misreads what was happening. The agent was not breaking the system so much as responding faithfully to what the system paid him for.
Consider each link in the chain. The panelist is paid to qualify, not to be truthful, and every screener is effectively a small exam where the passing grade is money, which is a kind of exam people learn to pass remarkably quickly. The recruiter is paid to fill quotas, so a recruiter who honestly delivers eight of the twelve requested respondents is, by every measure his agency tracks, a worse recruiter than one who delivers twelve with a few of us mixed in. The field agency is paid to deliver completed sample on a deadline, and auditing its own recruiters aggressively would mean missing deadlines and shrinking margins, so the audits stay gentle. And the brand, the only party in the entire chain with a real stake in the data being true, is also the party furthest from the point of collection, seeing nothing but the final report with its clean crosstabs and its confident recommendations.
Every layer is compensated on volume, and no layer is compensated on validity. When incentives and truth point in different directions at four consecutive links, what emerges at the end is not data with a bit of noise in it, but something closer to a manufactured product that has been shaped at every stage to look like data.
It is not just an anecdote, and it is not just an old one
You might hope that all of this is a story about focus groups in one city fifteen years ago, and that modern online panels, with their scale and their software and their attention checks, have engineered the problem away. The evidence, unfortunately, points fairly clearly in the other direction, and the useful thing is that some of the best evidence comes from researchers who had no commercial reason to sugar-coat what they found.
In 2020 the Pew Research Center published a large methodological study on opt-in online polls, running the same questionnaire across six online sources with more than sixty thousand interviews in total, in order to figure out how much of the data was coming from people who should not have been in the sample at all. They found that widely used opt-in sources contained “small but measurable shares of bogus respondents,” in the range of four to seven percent, compared with about one percent in panels recruited offline through residential addresses. More uncomfortably, these bogus respondents were not just noise; they tended to answer approvingly to almost anything, which introduced a small but systematic bias into the results. Eighty-four percent of them passed the trap questions designed specifically to catch them, and eighty-seven percent passed the too-fast-to-be-real checks. The standard quality filters, in other words, were finding almost nothing. (Pew Research, “Assessing the Risks to Online Polls From Bogus Respondents”, 2020)
Now, four to seven percent may not sound catastrophic on its own, and my point is not that panel data is fraudulent to the eyeballs. My point is that it took a research institution with a genuine methodological budget and sixty thousand interviews to detect it at all, which tells you what a typical commercial vendor, running attention checks on a fielding deadline, is probably catching. It also tells you something about the professional-respondent economy that many people in the industry talk about privately but rarely put in writing. Any environment where the same person can complete dozens of surveys a month across overlapping providers will create people who are extremely good at surveys, whose economic incentive is to keep qualifying, and whose relationship to the questions being asked is not what the researcher imagines.
There is a second, quieter problem sitting alongside the fraud one, which has nothing to do with dishonesty and which I think about more often than the first. It is the question of who is sitting on a panel in the first place. If you are a premium or even semi-premium brand, ask yourself honestly whether your actual customers are on a consumer panel. Would they have the time? Would the incentive mean much to them? People who are doing well tend to be busy and tend not to need the fifteen dollars, which is not a moral failing on anyone’s part, it is just how attention and money work.
A 2016 Pew study that evaluated nine online nonprobability samples found exactly this pattern in the data. The samples, the authors wrote, “disproportionately included adults without children, living alone, collecting unemployment benefits, and with low incomes,” which is more or less the profile you would expect of people for whom small survey incentives are worth the time. And the study’s most useful warning was about the standard reassurance the industry leans on: two of the least accurate samples were among the most demographically balanced on paper, because “matching marginal distributions doesn’t help if respondents within groups differ from their population counterparts.” In plain terms, your panel can match the census on age, income, and region and still be full of the wrong people, since the thirty-five-year-olds who joined a panel are not interchangeable with the thirty-five-year-olds who did not. Quota-balanced crosstabs make everyone in the room relax, and they should not. (Pew Research, “Evaluating Online Nonprobability Surveys”, 2016)
Even the fully honest, correctly recruited respondent does not quite give you her truest self, because answering questions is a social act and people manage how they come across, even to a form. Survey methodologists have documented the patterns for decades: satisficing, where you give the first acceptable answer rather than the accurate one; acquiescence, the polite tendency to agree; and social desirability, the big one, where people report the self they would like to be. Ask about gym attendance or vegetable consumption or how much advertising influences them and you will get the aspirational answer, reliably and in a predictable direction. Brand questions are especially exposed because brands are bound up with identity, and when someone tells a survey she would consider a premium brand, she is not lying exactly, she is describing the person she is on her better days.
Which brings me to brand lift
What I deal with now is far less intimate than a hotel conference room, mostly brand lift studies of one kind or another, and I can only imagine how the same misaligned incentives creep through at their own scale. The design sounds rigorous when you describe it: show ads to an exposed group, withhold them from a control group, survey both, and read the difference. But the respondents come from those same access panels, professionals and all. The response rates to in-banner and in-app surveys sit in the low single digits, which means the people who answer are a small and unusual minority willing to interrupt whatever they were doing for a questionnaire, and there is no particular reason to believe that minority looks the same in both arms. The questions are brand-perception questions, which is exactly the territory where people answer aspirationally. And respondents are not stupid; when a survey about insurance brands appears moments after an insurance ad, plenty of them can guess what is being tested and drift, helpfully or contrarily, toward an answer. In its own compressed way, it is my hotel conference room all over again, a room full of people performing the response they believe the occasion calls for.
What to actually do with this
None of this means that brand lift results, or survey research generally, are crap. I use these studies and I will keep using them. It means there is a nuance to hold onto and a set of incentives to keep in mind whenever a number arrives looking cleaner than the process that produced it.
The answer, as far as I can tell, lies in triangulation and in reading these studies directionally over time rather than absolutely in the moment. If your awareness or consideration numbers move gradually and coherently across many studies, seasons, and campaigns, that trend is telling you something real, because the distortions I have described are reasonably stable while your brand is what changes. If a single study shows a sudden, worrying shift, that is worth investigating. What is genuinely not worth agonizing over is why this campaign’s lift came in two points below the last campaign’s, because the honest error bars around any single reading are wider than the difference you are trying to interpret, and somewhere inside that difference are the panels, the professionals, the people who were charming enough to pass a vibe check, and everyone describing the person they are on their better days.
There is a companion problem on the other side of the measurement stack, which I want to write about separately. Lift tests on behavior, the rigorous end of the spectrum, have the opposite flaw: they measure something real with real precision, but only a narrow, short-term slice of it. Between the two essays the summary of advertising measurement comes out something like this: the rigorous methods see too little, and the broad methods have to be read with a knowing eye. I find that more freeing than depressing. Once you stop expecting any single number to be the truth, you can start using all of them for what they actually are, which is hints from a noisy world, some of them gathered over good food in a nice hotel by people who were mostly there for the iPod.