AI & technology
AI & data integrity (fraud)
27 indexed claims · 8 guests · peaked 2025-Q2

Overview
Survey fraud is no longer an anomaly — it is the ambient condition of online research. These podcast conversations map a crisis that has evolved from screener-gamers to coordinated global networks to AI agents, with a new normal where 30–40% bad data on a study passes unremarked. The dominant tension is structural: technical tools proliferate (23 fraud-detection vendors at last count) yet guests argue that no technology stack addresses the underlying incentive models rewarding volume over quality. A second tension runs through almost every conversation: AI is simultaneously the sharpest weapon in the fraudster's arsenal and the best available defense — an arms race no single actor can win unilaterally. The editorial center of gravity sits between resignation and urgency: the problem cannot be fully solved, but ignoring it will destroy the business model that funds market research.
From the corpus
- In a future of fluid, AI-mediated data access, the hardest problem will be distinguishing trustworthy signal from hallucination — and hallucination won't only mean LLM error; it will also mean comparing data that should never have been compared (different timeframes, contexts, countries) when easy access removes the friction that previously forced careful interrogation. — Lev Mazin, Ep. 24 ·
cl-lev-solo-020 - Detecting low-quality respondents requires building smart models that triangulate over many signals (the 'sum of all factors' or digital body language) rather than reaching for simple binary rules; data quality is a perpetually multifaceted problem and no blunt instrument will durably solve it. — Lev Mazin, Ep. 24 ·
cl-lev-solo-014 - The data quality problem resembles the antivirus industry's problem: the work is never done because new threats and bad actors emerge continuously, and AI is now both the attack vector and a source of defenses; the same technology that erodes data quality can be used to spot and improve it. — Shanon Adams, Ep. 49 ·
cl-50th-006 - Researchers should apply the same empathy to survey respondents that they apply to the consumers they are studying — translating questions into the respondent's language and generation, and remembering that some bad actors are economically desperate humans rather than fraud-bots. — Shanon Adams, Ep. 49 ·
cl-50th-007 - Crude defenses against AI-generated survey responses (e.g., blocking copy-paste) harm honest, careful respondents — non-native English speakers checking grammar in Word — more than they catch fraud; the right response is multi-signal triangulation, not blunt instruments. — Lev Mazin, Ep. 24 ·
cl-lev-solo-013 - Survey fraud has evolved from individual respondents gaming screeners, through co-located survey farms, to distributed global networks of people sharing attack techniques in real time on Discord and Telegram — each generation harder to detect and faster to scale. — Roddy Knowles, Ep. 8 ·
cl-roddy-knowles-004 - Companies should preregister their hypotheses and analysis methods before data collection, and on large secondary datasets they should split the data in half — mining one half and validating the conclusion on the other — to keep findings honest. — Ernest Baskin, Ep. 48 ·
cl-ernest-baskin-001 - Fraud detection is most powerful when passive behavioral monitoring — how a person actually interacts with the browser, how their responses are entered — is combined with active screener question responses; neither layer alone is sufficient. — Roddy Knowles, Ep. 8 ·
cl-roddy-knowles-007
Contributing guests
Episodes in AI & data integrity (fraud)
- aytm 50th Episode — Year-in-Review with Lev Mazin and Shanon Adams — Ep. 49 · Mar 3, 2026
- When the data is right but the decision is wrong with Ernest Baskin — Ep. 48 · Feb 24, 2026
- Scaling Consumer Insights: Design Thinking, Empathy, Tech & Analytics in Action — Ep. 24 · Sep 2, 2025
- The Real Cost of Bad Data: Survey Fraud, AI Agents, and the Fight for Data Integrity — Ep. 8 · May 13, 2025