Why quantitative understanding of effect sizes matters, even if all you care about is the presence of the effect

In reaction to my article with Andy King proposing post-publication review, Dan “Fast and Frugal” Goldstein writes: Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 6 days ago

What do we learn from bestseller regressions?

Gaurav Sood writes: I was reading ‘The Bestseller Code.’ The book reports results from some regressions of the form: bestseller or not ~ features of content This got me thinking about what you can recover from such an exercise. Say … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 7 days ago

Posterior predictive checking is for non-Bayesians too!

When I first started working on posterior predictive checking back in 1988, it was as a device for determining equivalent degrees of freedom for a chi-squared test for a model with constrained parameters–in that case, positivity restrictions in an image … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 8 days ago

“Over-coverage caught by pre-registration: 47 of 56 inside a stated 50% interval”

Alex Malinowski has a question about evaluating the calibration of interval forecasts: We publish interval forecasts under a pre-registration scheme: each forecast is serialised, hashed and timestamped into a Bitcoin block before publication, so the stated interval cannot be adju … | Continue reading


@statmodeling.stat.columbia.edu | 9 days ago

Eleven Kinds of Loneliness: Richard Yates and the tragedy of agency

Following up on Richard Yates (see last year’s post, Double Feature: Revolutionary Road and That Darned Chatbot), I came across his collection of short stories from 1962, Eleven Kinds of Loneliness. These stories are wonderful and deserve all the praise … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 9 days ago

Survey Statistics: equivalent models, equivalent weights (locally)

Last month we saw that the Times/Siena Poll is now using energy balancing weights (Huling & Mak, 2024). In a toy example, we saw under which outcome models these weighting methods might do well. I was inspired by Little 2004, who … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 10 days ago

Selection effects can go both ways (for taxi drivers as well as the rest of us)

You know how we talk about the two modes of microeconomic reasoning? For example, here: The logic of social science can work in two directions: generative modeling predicts behavior given assumed preferences, and inferential reasoning deduces preferences given observed behavior. … | Continue reading


@statmodeling.stat.columbia.edu | 10 days ago

He fit the same statistical models with three different software and got much different estimates. It’s another dimension of the multiverse.

Scott Cunningham writes: You’ll appreciate this I think. I ran Claude code on 96 specs for a popular difference-in-difference estimator with the identical specifications, ranging covariates only, for three languages (R, Python and Stata) and 2 packages for each. Different … Conti … | Continue reading


@statmodeling.stat.columbia.edu | 11 days ago

People sometimes talk about “the Jewish vote,” but what’s relevant is not really the Jewish vote or Jewish public opinion; it’s really about campaign contributions and the news media. Also similar with Mormons.

At the end of the second world war, Jews were a bit over 3% of the U.S. population, voted at a high rate, and were concentrated in the swing state of New York. Jews had two big issues–Israel and political … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 12 days ago

I don’t see journal review as a gatekeeping process that will keep erroneous articles from being published

Someone pointed to this post from last year, “If only Arxiv required researchers to sign at the top rather than the bottom of the page, none of this would’ve happened,” and asked about this statement of mine: “Seriously, though, setting … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 13 days ago

How generic language shapes the development of social thought

Recently in the sister blog: Generic language, that is, language that refers to a category as an abstract whole (e.g., ‘Girls like pink’) rather than specific individuals (e.g., ‘This girl likes pink’), is a common means by which children learn … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 14 days ago

Pretty maps of the NY mayoral election vote. (The meta-point here is that people have a (false) intuition that any complicated piece of information can be conveyed in a single plot.)

In reaction to my recent post, If Cuomo had been able to run against Mamdani head-to-head, would he have won?, sociologist Kieran Healy posted a pair of maps showing precinct-level results from the recent New York mayoral election. One of … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 15 days ago

Survey Statistics: poststratification without population level information

Poststratification uses population data on X to estimate E(Y) via E(E(Y | X, R = 1)), where R = 1 are survey respondents who provide Y and X. When the inner expectation “E” is estimated via Multilevel Regression, this is … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 17 days ago

Was this USDA survey really “redundant, costly, politicized, and extraneous”?

Joshua Brooks writes: I know you’ve posted on the topic more generally but don’t recall if you’ve discussed this in particular. Given the timing in relation to cuts in food assistance, It seems a particularly egregious example of the politicization … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 17 days ago

Rebecca Makkai points out: Fancy three-dimensional sets are easier to construct in books than in movies, but harder to explain

In a post entitled, “You’re Writing a Book. So Stop Writing a Movie,” Rebecca Makkai writes: You want to set your movie in a futuristic New York where every building has a flying car port on top and there are … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 18 days ago

It’s all about the Super Pacs: How the New York Times completely misreported campaign contributions in the Maine Senate race

Tom Ferguson came across this news article, Who Really Has the 2026 Midterms Cash Edge?, and was disappointed to see this completely wrong graph: The problem here is not the inclusion of no-longer-candidate Platner, as that’s noted in a footnote. … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 20 days ago

More scientists in the Epstein files, including a roboticist and an ESP researcher

I came across this webpage by Sheeva Azma entitled, “Here’s every scientist I have found in the Epstein Files so far.” She’s missing a few big fish: Dan Ariely (professor at MIT and Duke, Ted talk star, and teller of … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 20 days ago

Herman Chernoff

I recently learned from a blog comment that Herman Chernoff passed away last week at the age of 103. He was born the same year as my dad. I first met Chernoff–it’s not like he was a particularly formal guy, … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 21 days ago

Reviews of our Bayesian Workflow book from Bin Yu, David Spiegelhalter, Brad Efron, Christian Robert, Rohan Alexander, and Mine Doğucu!

Roughly speaking, Bayesian Workflow is to Bayesian Data Analysis in 2026 what Bayesian Data Analysis was to earlier Bayesian books in 1995: it builds upon everything that came before. With Bayesian Data Analysis, the big steps forward were: Going beyond … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 22 days ago

A ranked-choice election in Maine: Using voting data to understand preferences

Evan Rosenman writes: The implosion of Graham Platner’s Senate campaign in Maine has upended a marquee Senate race, leaving the state Democratic party just a few weeks to choose a substitute nominee. A planned nominating convention on July 25th has … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 23 days ago

Survey Statistics: quantifying uncertainty in ranked choice voting polls

We’ve talked about uncertainty in polls (see Margin of Error, Total Margin of Error, Total Margin of Error II) and we’ve talked about ranked data (see exploded logit !). A new paper, Rosenman & Liang 2026, looks at uncertainty in … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 24 days ago

“Making Statistics Work: Information Theory and Bayesian Inference”

I took a look at the above-titled book by economists Duncan Foley and Ellis Scharfenaker. It’s an interesting read, in many ways a throwback to the 1950s when a group of mathematicians brewed a heady mix of operations research, game … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 24 days ago

It’s all about the nonlinearity: An interesting statistical example of flaws in a voter impact index

The following came in the email the other day: I’m reaching out to introduce the Voter Impact Index, a new data tool from PowerMoves that assigns every U.S. zip code a voter impact score based on the recent competitiveness of … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 25 days ago

“Archaeology can’t give social scientists population or GDP, but here are some things we can measure that might be useful for social science.”

Apropos of our recent discussion on the estimation of historical population sizes, Sean Manning writes: Some archaeologists have measured house sizes for Gini-coefficient-style studies aside from studying human remains to measure nutrition and rates of illness. I think that was … … | Continue reading


@statmodeling.stat.columbia.edu | 26 days ago

“More bad science from JAMA”

In an abstract entitled, “Statistical dust and sweeping claims about maternal warmth,” John Richters and Everett Waters write: Alley and colleagues draw on mediation analyses of longitudinal data from Millennium Cohort Study to argue that their findings “highlight the critically … | Continue reading


@statmodeling.stat.columbia.edu | 27 days ago

18 Associate Editors resign from Statistics and Computing editorial board: Problems with commercial scholarly publishing, and what does this all mean?

I was cc-ed on a message sent by 18 members of the board of the journal Statistics and Computing, quitting their posts because the publisher (Springer) has announced a new policy whereby all authors will have to pay publication charges. … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 28 days ago

“A medical journal says the case reports it has published for 25 years are, in fact, fiction”

Retraction Watch reports: A Canadian journal has issued corrections on 138 case reports it published over the last 25 years to add a disclaimer: The cases described are fictional. Paediatrics & Child Health, the journal of the Canadian Paediatric Society, … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 29 days ago

Is fabricating data worse than fabricating results? Is failing to correct a known false report more or less serious than making the false report in the first place?

Andy King writes: I have a question for you–and, if you think it worthwhile, for your readers. A few weeks ago, I was deposed by Harvard’s lawyers in the lawsuit between Francesca Gino and Harvard. Much of the questioning focused … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Survey Statistics: toy example for energy balancing weights

Last week we talked about The Big Changes Coming to the Times/Siena Poll: New weighting variable: support score = E(2024 vote | other X variables). New weighting method: energy balancing (Huling & Mak, 2024) Ben Schneider helpfully blogged about energy balancing … Continue readin … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Claude builds 3D Hamiltonian Monte Carlo animation in one shot with anaglyphs

This post is from Bob The sausage So as not to bury the lead (or “lede” if you want a mid-20th-century newspaper vibe), check out the this 3D HMC animation generator. It can render regular animations or produce anaglyph 3D … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

A message for Carol Tavris

Dear Dr. Tavris: I saw in a recent issue of the Times Literary Supplement that you have been critical of the “chambermaid” study which purported to show that people were losing weight without changing their diet or exercise. I agree … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Turning chaotic sensitivity from a bug into a feature: Using physical modeling and deep learning to alter the paths of storms and mitigate extreme weather events

Qin Huang, Moyan Liu, and Upmanu Lall write: Extreme weather events, e.g., droughts, floods, heatwaves, and freezes, are increasing in frequency and intensity, posing severe socio-economic impacts as growing populations heighten exposure to risks that conventional infrastructure … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

The NIH wants to “Measure and Reward Scientific Impact and Replicable Research Practices.” Here’s my recommendation to the NIH director: you can start by no longer suppressing government reports whose conclusions happen to not be in accord with your ideological preferences.

This came in the email from the U.S. National Institutes of Health: How Would You Measure and Reward Scientific Impact and Replicable Research Practices? As NIH continues efforts to strengthen rigor, reproducibility, and public trust in science, we are seeking … Continue reading … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

2015-vintage replication-crisis-era junk science floats into the news

So, I came across this news article titled, “Riley Thinks Suits Make the Coach. Research Says He Might Be Right.”: The suit had a classic name: the Clark Gable. Navy blue and cut just right, it was the creation of … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

A new episode in the Francesca Gino case

Andy King writes: 𝗪𝗵𝘆 𝗛𝗮𝗿𝘃𝗮𝗿𝗱’𝘀 𝗹𝗮𝘄𝘆𝗲𝗿𝘀 𝘀𝘂𝗯𝗽𝗼𝗲𝗻𝗮𝗲𝗱  … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

The high cost of split R-hat

This post is by Bob. I’ve been thinking a lot lately about R-hat given that I’m using it for online converging monitoring in our new Walnuts implementation. In that setting, where I use Welford accumulators to update R-hat estimates every … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Guess who’s getting the big-money donations in the Maine U.S. Senate race?

Just in time for July 4th, Tom Ferguson, Paul Jorgensen, Matthias Lalisse, and Jie Chen share the above graph and write: What can one Senate race reveal about the hidden machinery of American politics? In Maine, donor patterns expose how … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

The optimizer’s curse

The above sketch shows a decision tree. The circles are uncertainty nodes and the squares are decision nodes. Read the tree from left to right: to start, there is uncertainty of which of the strata i=1,…,I you will be in. … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Survey Statistics: Big Changes in the Times/Siena Poll

Yesterday Nate Cohn wrote about The Big Changes Coming to the Times/Siena Poll, with more details in their poll of Maine. Say we want to estimate average Platner support in Maine’s likely electorate, E(Y). But we only have survey respondents, … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

OK, I guess Lawrence “Epstein” Krauss didn’t follow his brother’s advice.

The former Arizona State University physicist reported in 2018 this advice from his “religious right wing law professor brother” [that’s Krauss’s description, not mine]: Therefore i think you should pursue a mixed strategy. On the one hand, you should non-aggressively, … Continue … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Cheapskate evolutionary biologist underpays his statistical help

OK, this one was funny. I searched the Epstein files for “statistician” and found this receipt from biologist Robert Trivers: Only $1000 for the statistician??? What a cheapskate! Especially given that he said the statistician “did an outstanding job.” Given … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

The Anthropic Principle in Statistics and Science (my talk this Mon 29 June, 4:20pm London time)

The Anthropic Principle in Statistics and Science The anthropic principle in physics states that our existence implies certain constraints on the natural conditions under which we evolved. In statistics, a corresponding anthropic principle can be used to infer properties of … Con … | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Bayesian Workflow exists as a physical book!

We’re very excited about this book. It’s the result of several years of effort. You can order from the publisher or from Amazon. Here’s the book’s webpage, which includes the data and code for the book’s examples and case studies, … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Out of the frying pan and into the fire: Scientific American returned to form, and then this happened:

Last month I wrote the following post. I scheduled it for November, but then some Scientific American-related news arose, so I’m bumping it up in the schedule. First, here’s my post from May: I’m not saying this is the same … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

“Springer Nature has removed two studies by Max Planck.”

Jim Moody points to this news article, “Why have papers by one of history’s most famous physicists been retracted? Springer Nature has removed two studies by Max Planck. A bot may be to blame.” If you’re gonna retract something from … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Supplement that alphabetized display with another graph showing the states in a more informative order.

I just wrote a long post inspired by a recent post from economist Paul Krugman. Krugman’s post was good, but I’m annoyed that his graph (reproduced above) lists the states alphabetically. Don’t do that! It’s called the Alabama first error. … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Structural equation modeling (SEM) and positive definiteness

This post is from Bob. Mitzi and I were swotting up on structural equation models (SEM) for our class this past Monday at the Modern Modeling and Methods (M3) conference at Fordham University. It was a lot of fun and … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago

Getting justice can require a lot of effort, and usually at some point we’ll just give up, which is what the cheaters rely on.

I just read this compelling op-ed by Brendan Ballou, “One Man Stole $660 Million. He’ll Never Pay It Back,” which tells the story of several brazen white-collar criminals who avoided prosecution for federal crimes by the simple expedient of bribing … Continue reading → | Continue reading


@statmodeling.stat.columbia.edu | 1 month ago