Practical guidance on study design, analysis, and evaluation.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}
Rates range from a few hundred to several thousand dollars depending on scope, not on how complicated your data feels to you. Here's what actually drives the price, and how to get a quote you can trust.
Most researchers call a statistician after the data is collected. By then, the decisions that would have mattered most — sample size, measurement, randomization — are already locked in.
A reviewer questioning your model isn't automatically a rejection. It's a request for evidence — and how you respond matters more than whether the original model was perfect.
Randomizing schools, clinics, or teams instead of individuals means your effective sample size is much smaller than your headcount suggests — and a power analysis that ignores this will underestimate what you need.
Funders don't just want to know your program happened. They want evidence it caused the outcome you're claiming — and evaluations that skip that distinction tend to lose credibility exactly when it matters.
A single accuracy number can hide the exact failures that matter most — subgroup gaps, instability across runs, and confident wrong answers. Here's what a more defensible evaluation actually measures.
Most analysis problems that look like modeling problems are actually data problems nobody checked for first. A short audit up front catches most of them before they cost you time.
Both handle missing data under the same core assumption, but they work differently and fit different situations. Here's how to decide which one your analysis actually needs.
A statistical analysis plan written before you see the data is one of the cheapest ways to protect a study's credibility — and reviewers, funders, and IRBs increasingly expect to see one.
Published trials made a class of antidepressants look almost uniformly effective. The FDA had the trials that never got published, and the real picture was a coin flip. A case study in publication bias, and why the literature you can read is not the literature that was run.
An automated essay-grading engine gave top marks to student essays that made no sense at all, as long as they used long sentences and fancy words. A real case study in Goodhart's Law, and why optimizing a proxy metric isn't the same as measuring the thing you actually care about.
A machine-learning model hit strong accuracy telling wolves from huskies — until researchers checked what it was actually looking at. A real case study in shortcut learning, and why a good accuracy number can hide exactly how a model is cheating.
A dead fish, put in a brain scanner, appeared to respond to photos of people with statistically significant brain activity. It was a joke with a serious point — about what happens when you run thousands of tests at once and only report the ones that worked.
Give a room full of physicians a positive screening test and one prevalence number, and most will badly overestimate the odds the patient is actually sick. A short primer on base-rate neglect, and why 'accurate' tests aren't the same as trustworthy results.
The Air Force wanted to armor the parts of returning bombers riddled with bullet holes. A statistician told them to armor the parts that had none. A short primer on survivorship bias, and why the data you don't have can matter more than the data you do.
Berkeley's 1973 admissions numbers looked like textbook discrimination against women — until anyone checked department by department. A short primer on Simpson's Paradox, and why the number you aggregate is a choice, not a neutral fact.
Scared Straight felt like it worked: kids were shaken, parents saw a difference, the numbers looked good. The best evidence says it made things worse. Here's why a good story keeps fooling careful people, and what to ask for instead.
{{ p.excerpt }}
No posts tagged "{{ activeTag }}" yet.