← All episodes

Ep. 68 · Jul 21, 2026 · Q3 2026

Data quality is not a filter, it's a discipline

Hosted by Molly Strawn-Carreño & Stephanie Vance · aytm

Molly Strawn-Carreño

Listen

Watch on YouTube · Apple Podcasts · Spotify

Summary

Hosts Molly and Stephanie discuss why data quality is a business issue that affects stakeholder trust and decision speed, introduce a four-P framework (prevent, protect, purify, prove) for moving from filter-first to discipline-first data quality, and outline red flags to watch for in vendor evaluations.

Your hosts

Molly Strawn-Carreño — Host, aytm
Stephanie Vance — Host, aytm

From this episode — top claims

  • The vision of AI systems that constantly learn and stay connected to one another already exists in human form — those connected, always-learning intelligences are people. — Ep. 49, Molly Strawn-Carreño
  • The right way to read a recurring say-do discrepancy is as signal rather than noise — a reframing of the gap not as a problem to close but as information about what creates it. — Ep. 49, Molly Strawn-Carreño
  • Storytelling in insights is not a mysterious creative craft; it is simply connecting with another human being to convey information — a skill insights professionals already possess but fail to recognize in themselves. — Ep. 49, Molly Strawn-Carreño

Full transcript

Auto-generated captions — speaker labels aren't always available and wording may be approximate.

0:00 The first layer that we're talking about here is prevent, which is about the survey instrument itself. And the question to ask is is the structure of the survey working for me or is it working against me? So, for example, a 20-minute survey on a phone is just a fatigue factory. And fatigue is actually the largest single source of unattentive responses, which is bigger than fraud. These are respondents who are genuinely trying to give their feedback. You've just made it a little bit too difficult for them to do so.

0:30 So, looking at that structure and making sure that you've set up an environment for well-meaning respondents to find success. Hello, fellow insight seekers. I'm your host, Molly, and welcome to the Curiosity Current. We're so glad to have you here. And I'm your host, Stephanie. We're here to dive into the fast-moving waters of market research, where curiosity isn't just encouraged, it's essential. Each episode will explore what's shaping the world of consumer behavior, from fresh trends and new tech to the stories behind the data.

1:03 From bold innovations to the human quirks that move markets, we'll explore how curiosity fuels smarter research and sharper insights. So, whether you're deep into the data or just here for the fun of discovery, grab your life vest and join us as we ride the Curiosity Current. Data quality has been acclaimed for a long time. Every vendor says their data is clean. Every platform says its respondents are real.

1:31 But what if you could actually measure that? Today, we're talking about what it looks like when an industry stops asserting quality and starts proving it. So, today, no guest. It's just Stephanie and me. We've been spending a lot of time lately thinking about data quality, not just as a feature that vendors can check off the list, but as a real business issue, one that shapes decisions, timelines, stakeholder trust, and whether insights teams are going to get a seat at the table. We wanted to talk through a framework that's been shaping how we at AYTM think about this.

2:03 Yeah, and there's a bigger shift happening right now that I think is really worth unpacking. This conversation, honestly, has been running on claims for years, and that's starting to change. And so, I'm genuinely excited for us to get into it today, Molly. Me, too. And I'm just going to start right off with saying something that I feel researchers know intuitively, but not necessarily something that is said out loud or comes out in every conversation, which is the cost of bad data isn't just the cleaning bill to fix it. That's actually probably the smallest part of that.

2:35 Absolutely. I love that. Let's start there, Molly. Unpack what like some of those costs are for us. Yeah, I I feel that there's four costs that doesn't exactly show up on an invoice, per se, but is still there and is is weighted as a cost, which is First of all, wrong decisions at scale, reruns and rework that you have to pay twice the money for, and it often takes up to three times as long to get that done. Stakeholders' trust is not cheap to lose, and it's very expensive to rebuild, and decision velocity. So, when teams don't trust the data, they ask for more, and that whole organization slows down. And I'm just going to sort of leave it at that cuz I have some thoughts on this, but what do you think, Stephanie, is from your experience in in the seat with clients and in the seat with researchers, what is the most impactful and costliest in your experience in and in practice?

3:30 First of all, reasonable people can disagree with that. I want to start by saying that it's easy and tempting, for me anyway, to want to say wrong decisions, because like that's at the heart of what we're doing. In reality, it's the erosion of trust, whether that's between the supplier and the client, which is where I have historically felt it as a supplier-side researcher, or between that brand-side researcher and their internal stakeholders. Beyond that, I think there's also this compounding effect that we should talk about. Each of these costs makes the others worse. Wrong decisions erode stakeholder trust.

4:06 Eroded trust slows decision velocity. Slower velocity creates pressure to run more studies. That creates more opportunities for the same problems, and we're in this data quality loop of hell. Cycle, yeah. Circle of hell. That behavior change and how it compounds off of one another, and how it changes the way that the workflows work at the organization level, not just the individual users level. What actually is the issue and never shows up in a postmortem, is never shown up as this is the core problem that's fueling and spinning off all these other issues.

4:41 Absolutely. That that postmortem is usually this very band-aid solution of a few people showing up to the table, a couple getting their hands slapped, and then we're at a change in process, and then we're moving on from there. But in reality, to your point, it is a cost that everyone pays. The insights orgs lose trust. The brand manager loses confidence, and the business loses that decision velocity. So, it's critical. It's critical work. There is a phrase that really resonates with me in AYTM's data quality philosophy, and it goes like this, "Data quality is not a filter, it's a discipline."

5:18 Those are not the same thing, and I think most of the industry has been optimizing for the wrong one. Molly, to get us started through kind of thinking through this, walk us through what filter-first thinking looks like. Yeah, so filter-first thinking is I feel like a lot of what the industry benchmark is that treats data quality as something that has to happen up front, but is not something that is part of the full process. It's sort of just accepting that contamination, right?

5:45 It's accepting that bad data exists and it is what it is, and then just try to remove it on the back end. But, the issues that we're finding is a lot of the research on research from our industry is showing that 70% of fraud actually slips through those standard methods of cleaning. Um and that's always changing, and fraudsters are always getting savvier and more innovative. So, the job that you're doing at the front is actually just a small fraction of what it actually should be. It It's It's a completely different mindset. The practical consequence of just thinking with filter-first mindset is that your cleaned end size is not necessarily exactly what you think it is. The number on the report has already been filtered.

6:24 But, again, like I said, that filter has missed most of what should have been caught and still creates issues. Yeah, we chalk it up as noise, and I love the phrase that you said at the beginning, which is accept contamination. And I think that that is just reflective of the fact that we think we have to just accept this. It's just like if you monthly But, it's just the industry norm. It's just the cost of doing business. Exactly. Yeah. Whereas, conversely, I think discipline-first means building the conditions where good data is the norm and where it can thrive, starting before the survey ever goes live, before any respondent see a question.

7:05 The cost to run either of these approaches, and this is the part that's interesting to me, it's really the same. It's just that the outcomes are wildly different. So, that's a That's a great point, Stephanie, about how it actually needs to be more of a holistic process. But, what does it actually look like? Again, I'm going to ask you in practice questions, because I feel like that's what I would want to know, super valuable as a practitioner. What does it look like when a research team has actually internalized the discipline-first approach? And what's the difference, perhaps, in how they set up studies? How do they choose vendors?

7:38 How do they present results? All those kind of good things. Absolutely. Well, I I think the main thing about this shift is it changes where in the research life cycle you invest attention and action related to data quality. Filter first, to your point, puts weight on post-collection processes. Discipline first puts weight on design. It puts it on panel selection, instrument quality, earlier in the process.

8:06 That's a shift that I'm just firmly convinced we need to see across the industry. And I would turn a question back to you, Molly. What do you think it takes for the industry to move toward discipline first as the default? Is it a buyer change? Is it a vendor behavior change? Is it both? I think the answer is actually very simple, which it's just what does the market ask for? It's the market dynamics. It's the economics behind it.

8:34 Clients really have to demand this and be making vendor selections, making choices about what types of research that they're using and not using and in what ways. They have to demand this in order for vendors to truly feel the pressure to respond. There are absolutely vendors out there that are doing this because it's the best way to present data. Ethically, we have to present the best data that we can and so we're going to adopt this mindset, but there does need to be that external push to manage change in order to say, "We have to make this investment. We have to make this change. It can't be business as usual because we're feeling that pressure to change."

9:11 Yeah, and I'll just add to that, you know, I have never met a supplier-side researcher who didn't care about data quality. That's not what's going on. I think where that could go out the window, though, is when it's made into a competing priority with cost or speed. So, I agree with you that change in demand from the brand side is key because it allows that priority to surface as something that we get to, that we need to pay attention to.

9:39 Yeah, I mean, the life of a researcher is hard. You have so many competing things that you're balancing, and this type of mindset takes time, it takes effort, and it takes workflow change. And then you have to communicate that change, right? To your clients and your stakeholders. This is going to take more time, or this is this is going to be a different approach to this. And having to manage that conversation is is also not easy. Right, let's revisit this instrument as something when a client is done and ready to hand it off for programming, that's not always a conversation they're ready to have at that point. But, my goodness, better to have it then 2 weeks later when we're analyzing data and it doesn't make a lot of sense to us.

10:18 So, I want to talk a bit more about the framework that we've talked about a bunch. And what I find useful about the framework that we've been working with here at at AYTM is it gives a specific place in the process where you can look to if something goes wrong, or there's a challenge that you're facing. And when you're evaluating whether it's something that will go wrong. So, the four layers are prevent, protect, purify, and prove. And we heard this a little bit because our head of data quality here at AYTM, Jonathan Goodbread, recently spoke at the 2026 Quirk's Virtual Session, Ensuring Data Quality, Security, and Ethics, which was discussions around data best practices for better research.

11:00 He introduced this idea about the the 4P and the 4P framework, which I think is really helpful for vendor evaluation as a whole. So, let's just dive right into it. The first layer that we're talking about here is prevent, which is about the survey instrument itself. And the question to ask is, is the structure of the survey working for me, or is it working against me? So, for example, a 20-minute survey on a phone is just a fatigue factory. And and fatigue is actually the largest single source of unattentive responses, which is bigger than fraud. You know, these are respondents who are genuinely trying to give their feedback. You've just made it a little bit too difficult for them to do so. So, looking at that structure and making sure that you've set up an environment for well-meaning respondents to find success. Uh the second one, protect, is about the sample. Uh the question you should be

11:57 asking is everyone who's in this actually really who they say that they are because every fraudulent respondent, you can actually think of it actually costs twice, right? It costs once to pay for when you were thinking that they were actually there and well-meaning and they they were going to give you some good data. And the second cost is to clean the data when you take them out. So, the goal is not necessarily to block them and remove them after the study cuz you've already incurred the cost of having them in your study. The point is to protect them uh from entering your study at all. Preventing them from coming into that study at all.

12:32 Absolutely. And then to jump into that next P, purify to me is where 2026 is genuinely different from even 2 years ago. The threats have evolved. AI-generated open ends, response farms, which are real humans, coordinated sharing infrastructure, and the attention ceiling that you were kind of referencing, Molly, that that good design can reduce but it cannot eliminate. If a vendor's infield strategy only addresses one of these three, it's solving at best a third of the problem.

13:04 Yeah, it's not going to work. It's not. And then finally, the proof layer is philosophically the most important to me. It's the difference between our quality is good, which is a claim, and something concrete that a CMO can actually read, compare, benchmark, and hold you accountable to. So, what's interesting is that you said that proof is the most important layer philosophically in the process. However, it's the one that can get skipped the most often, which is really interesting to me because it's the one that, you know, is going to be the most defensible. It's the one that matters when a stakeholder like your CMO challenges your numbers or challenges any of the takeaways here. Without a per study quality record, you you you can't defend anything. You you have no leg to stand on when they start asking those very specific questions. So, without that comparability across studies, you can't tell if it's your program is getting better or if it's just getting

14:04 bigger. You're just doing more things. Right. And you know, one of the things that these four P's holistically do is that they allow you to sort of assess what we call a data quality report at the study level for every study. And it raises the question for me, what would it change? What would it look like for an insights team if they assess data quality? They read that data quality report for every study before they even looked at the results every single time.

14:33 And that is the action that John, our head of data quality at AYTM, kind of closes his talk around. I think it's harder than it sounds, but I think this is exactly where as an industry we need to be heading, as a discipline we need to be cultivating. Well, let's get specific here. Let's just say I'm a senior insights director and I'm sitting across from a research vendor I I or evaluating one. I'm in the evaluation process.

15:00 What are the questions that I should be asking when it comes to data quality? And maybe even more important, what are the types of answers that I should just refuse? Like what's a red flag that I should be out on the lookout for? Okay, well, first off, I think we have to be asking suppliers whether their fraud detection is single signal or relationship based. Um, single signal catches the obvious stuff, but relation-based detection catches coordinated behavior, and that's going to hit your response farms, things like device sharing, proxy patterns.

15:33 If they can only describe one signal, your vendor I mean, they're not seeing the full picture, as we've talked about. Also, you got to ask about what the in-field strategy looks like across AI-generated responses, across response farms, and across inattentive respondents distinctly. If they only talk about one of those categories, again, they are solving for about 1/3 of the problem. And these questions and what we're getting at here is there's a broader signal in in this and this is a moment that's worth naming. For example, AYTM is moving to towards publishing data quality figures publicly about all of our different surveys that we have running, and it's it's the first time that anyone in the industry has put out the the actual auditable numbers out in the open. And it changes the conversation from just our quality is good to actually here are the numbers and here is the methodology and here is how you can better audit us. It begs the

16:32 question of what happens to the industry then if this transparency completely becomes the norm? Does it raise the floor for everyone or perhaps it maybe only benefits the vendors who are already investing in data quality? All the business will just go to them if everybody starts talking about this or will it challenge everybody to to step up? Yeah, I think that's a great question and I will be the first person to say I don't have a crystal ball.

17:00 Why not, Stephanie? I don't, but but the answer to refuse as an evaluator of vendors of of panel is trust us. That's the answer to refuse. Transparency is becoming the buying criterion. Vendors who can show their work will win. Those who can't should be pressed until they can. Yeah, and I think my answer to that is is also that the credibility gap will widen as people are being more challenged in this. And insights teams that can't defend their numbers and be transparent are going to become indispensable, whereas the ones who can't are going to be asked consistently to rerun a study until they may be replaced by a different vendor. And and that truly is the business case for caring about all of this.

17:46 Absolutely. So, we've talked a lot about practical applications, but let's talk about perhaps what an insights professional can do tomorrow. So, Stephanie, what is a question that a researcher should ask their vendor this week? A new vendor or one they're evaluating or perhaps a vendor that they have an ongoing relationship with. That's maybe something they've never asked before. What is that question?

18:12 I think that the question that I would suggest that brand-side researchers ask their vendors or supplier side who use external panel is to ask the question, "Walk me through your data quality framework." I want people to forget nitty-gritty details. I don't want to hear about panel books, cleaning rates, source blends. Ask about the framework. If they are not being proactive and driven by a strong point of view, they're scrubbing data, and that's a band-aid. And to the whole point of this conversation, that is just not enough anymore.

18:48 Yeah, and I feel like that if it's checking a box, you'll be able to get a feel by their answer in this that they're they're just ticking a box and moving on with their life and sending you what they have, or they're actually driven and passionate and interested in staying ahead of these topics and that they have an internalized way that they go about this process and it's not just something that they do and then move on. I feel like this answer, I mean, even what they tell you, but also the thematic behind it, their approach to how they answer it is also going to be telling as well.

19:21 Absolutely. Well, thanks so much for joining us today, as always, on the Curiosity Current. If this conversation sparked something new for you, we'd love for you to subscribe and leave us a review. It really does help people find our show. And this conversation about data quality is not over. We will be back with more episodes, with more experts, with deeper points of view. So, stay tuned for that. And until next time, we encourage you to keep asking the questions that matter, and we'll see you next time on the current.

19:52 The Curiosity Current is brought to you by AYTM. To find out how AYTM helps brands connect with consumers and bring insights to life, visit aytm.com. And to make sure you never miss an episode, subscribe to the Curiosity Current on Apple, Spotify, YouTube, or wherever you get your podcasts. Thanks for joining us, and we'll see you next time.

Produced by aytm

Curiosity Current is made by aytm, the consumer-insights company. Guests speak for themselves; the synthesis is ours. About aytm