↵ select ↓ ↑ navigate esc close

Goldilocks questions

Unpublishable Papers ·

Portrait of the scientist as a young burglar

Running up to the launch of GPT-5, Sam Altman claimed that his company’s new model would display “PhD-level intelligence.” Lots of people, particularly PhD students, were annoyed by his declaration. They replied that there is no innate measure of intelligence; that getting a PhD requires not computational power, but hard work and curiosity; that if a system can’t even label a map, it doesn’t matter if it can write a grammatically-sound dissertation in fifteen minutes or less.

Not to mention that GPT-5 mostly failed to correct the hallucinations, basic fallacies, and other errors displayed by prior iterations. PhD-level intelligence? Start with fifth grade first!

Subscribe now

And so on. I don’t particularly care about the validity of “PhD-level intelligence” as a metric for the performance of large language models. Altman used the term as a marketing ploy, not as a quantifiable metric that GPT-5 had actually hit.

But I was struck by all the confusion over what, in principle, PhD-level intelligence could even mean: that is, what kinds of abilities are ostensibly cultivated during a doctoral degree that a layman does not possess.

Throughout the first year of my PhD program, I’ve wondered this quite a lot. What am I learning, exactly, for 8-10 hours a day? Maybe some new techniques — basic programming, statistics, whatever — but I could have learned all of that as an undergraduate. Am I gaining unique expertise? By the time I’m done, I will definitely know way more about certain scientific subfields than most people on Earth. But you can also cultivate expertise by working, say, as a labor lawyer or a chef.

Nor is any of this a matter of raw intelligence, such that a PhD simply acknowledges preexisting ability, like a more prestigious version of Mensa. I know lots of smart graduate students, but I can also think of several people smarter than me who chose career paths that I am probably not capable of following.

So if the difference between “Eli Elster” and “Dr. Eli Elster, PhD,” is not technical skill, or expertise, or intelligence . . . what is it? Theoretically, what is PhD-level intelligence supposed to mean — or in other words, what is supposed to be different about someone who has acquired a PhD, versus someone who has not?

Here’s my answer. A well-designed PhD program teaches you how to ask good questions. Trivial as this may sound, no ability is more central to scientific progress than asking good questions. And doing it is much, much harder than you think.

Why is that important?

Last October, I went for a walk with my advisor, Manvir, to talk about some ideas. I rambled for ten minutes about some potential study on the evolution of storytelling; honestly, it wasn’t worth remembering. Manvir listened quietly, nodding and staring at his flip-flops as we looped aimlessly around Davis.

When I finally petered off, he said: “Okay, okay. Word. But why is that important?”

You may have heard that there is no such thing as a stupid question. In principle, this is true, because unadulterated curiosity is beautiful. In practice — that is, if you’re trying to become a scientist — it is false. Rather than ‘stupid,’ though, it is more accurate (and kinder) to say that some questions are either not worth pursuing or not pursuable in the first place.

For instance, consider a question like: “Do more open-minded people prefer snowboarding over skiing?” You could totally answer that question! Just ask people to take the OCEAN personality test, then answer questions about their snow sport preferences. Cool. Now do various statistical tests to see if there are any correlations between open-mindedness and snowboarding, and voila, you have an answer to your scientific question.

Top 10 Most Gnarly Snowboarding Games | Articles on WatchMojo.com

Gnarly, bro! Hella open-minded, potentially, contingent on p-value!

Great! But why do we care about that result?

Intuitively, we don’t. The fact that open-mindedness is correlated with snowboarding over skiing, or vice versa, is a fun fact that you see on social media and forget about four seconds later. You could publish the result and no one would ever cite it or notice it or think about it ever again.

Here’s another question: why do humans make art? Unlike the first question, this seems worth pursuing — aesthetic practice is a strange and central feature of our existence, and everyone would like to know why we do it.

But there’s a problem: this question is not specific enough. It is hard, perhaps impossible, to come up with a way to systematically answer it in its current form. So this too is not a good scientific query.

From these two examples, we can (by way of omission) identify a few characteristics of good scientific questions:

Pursuability

  • Can we identify a method capable of producing an answer that would seem to answer the question?

Generativity

  • Does answering the question reveal new research directions, either by applying new methods or opening new gaps in the literature?

Fundamentality

  • Does an answer to the question alter or improve our understanding of many other scientific domains?

The relationship between snowboarding/skiing and open-mindedness is pursuable, but not generative or fundamental — the answer doesn’t seem to matter for other fields, and it doesn’t suggest obvious next steps for research. Conversely, modeling the evolution of art is generative and fundamental, but not really pursuable per se — the answer doesn’t seem to follow from an actionable method, because the question is too big to handle.

So our two examples are both bad questions, but for very different reasons. One is too big, and one is too small.

We must instead try to pose Goldilocks questions. A Goldilocks question, of course, is one that is just the right size. It isn’t so big that you can’t answer it, but not so small that it isn’t worth answering in the first place.

Notably, once you’ve posed such a question correctly, the answer often falls right out of the formulation. Not literally — you’ll still need to collect the data, apply the methods, and so on. But that’s often just a matter of doing lots of grunt-level work: hence why research assistants are regularly hired to assist or take over the process.

As Einstein put it: “If I had an hour to solve a problem and my life depended on the solution, I would spend the first 55 minutes determining the proper question to ask… for once I know the proper question, I could solve the problem in less than five minutes.”

Albert Einstein | World-famous theoretical physicist | New Scientist

This article is bullshit. Great science is all about smoking tobacco (see above).

Two paths to finding the porridge

Coming up with Goldilocks questions might seem like a means to an end, rather than an end itself. It isn’t. My sense is that doctoral students often spend years developing the ability to ask these kinds of questions, if they manage to do it at all.

Sometimes they’ve gotten there before starting their program, which makes it easy for them to complete their dissertation without hiccups. Good for those people (he said, resentfully). But this seems to be rare. Usually, students must follow one of two paths before they can figure out how to ask the right kinds of questions.

Small → big → just right

The first path is small → big → just right. Consider a typical student who follows this path. Let’s call him Dennis.

Before starting his doctoral program, Dennis worked in a lab where he was tasked with carrying out a pretty niche task: say, coding responses to surveys about attitudes toward violence. Methodologically, he knows what’s up. But his role in the lab was oriented more toward assisting senior scientists rather than developing his own projects. He also never read very widely outside his field.

So upon arriving in graduate school, Dennis can come up with lots of small questions — he knows how to copy the methodological structure of projects he’s worked on before. He tells his advisor he wants to study the relationship between personality traits and hobbies. But then his advisor asks him: Hey Dennis, why are these questions important?

Dennis realizes he isn’t sure. At night, he starts to wonder whether he possesses the basic curiosity one needs to come up with interesting questions, or if he’s just a research robot unblinkingly following a track that he, in fact, doesn’t care all that much about.

Mental Health of College Students Is Getting Worse | The Brink | Boston  University

Woe is Dennis.

If all goes well, though, Dennis will start to read widely. He’ll have lots of conversations with his mentors and his peers. Ultimately, he’ll identify the basic interests that inclined him toward his field in the first place. He never really cared about personality trait/hobby dynamics; what interests him, he realizes, is the way that personality traits change over our lifespan!

From that point, he can apply his methodological knowledge to narrow his well-identified interests into Goldilocks questions. For instance: can open-mindedness continue to shift as we age, or is it concretized after some point in development?

Big → small → just right

The second path is big → small → just right, embodied by a student named Kat.

As an undergraduate, Kat bounced between many different majors, read anything and everything, and collected an intriguing but disparate CV. Her research experience was limited. When she applied to graduate programs, her application contained several grand areas of interest and few concrete ideas.

Upon starting her program, however, Kat realizes that she doesn’t really know how to create projects within those areas of interest. She simply has the interests themselves. She likes books about storytelling, and music, and psychedelics1 — but she doesn’t really know how to make a publishable study about any of those things. Interests are not the same as questions. And without questions to answer, she will never be a scientist.

Despair! Kat wonders about her career choice. Maybe she should’ve been a journalist instead.

But she doesn’t quit. She starts to keep a big list of ideas, and starts reading papers that are not small (never!) but that explore her interests in ways that still strike her as novel and ambitious. Most of her ideas are bad. Over time, though, she learns to copy the way that her intellectual forebears have asked questions, and her ideas get better. She also talks to her peers — mostly people who are following the small → big → just right path2 — and uses their framings to guide her own.

She shares these ideas with her advisor, and her advisor agrees that most of them won’t really work. But that’s okay, because a few of them definitely will. She hears a Hemingway quote one day: I write 99 pages of shit for every one page of gold.

Teach the children well

These are both legitimate paths to learning how to ask Goldilocks questions. The truth, though, is that many PhD students probably never figure it out.

In some cases, that might be the fault of their work ethic or intelligence. But usually, it’s because their advisor doesn’t understand what it is that their students are supposed to be learning. They push their advisees to pursue projects they didn’t come up with on their own, or suggest that they focus on teaching, or only make time to discuss action items. In short, they emphasize everything but the cultivation of Goldilocks questions.

This is bad practice, and it has predictably bad results. A big → small → just right student can start to feel like they’ll never focus enough to publish a paper. So they quit their program, or switch to asking small questions out of desperation. A small → big → just right student can stick with it more easily by publishing on all manner of tiny and useless questions; but that means they too cannot become a worthwhile scientist, because no one is teaching them how to be one.

I suspect that I’ve gotten lucky — my advisor is uniquely focused on helping me figure out how to ask Goldilocks questions, and I think I have benefited tremendously from it. 3

Time Bandits | The New Yorker

Albert Einstein and Kurt Gödel on one of their famous walks.

But his focus seems to be exceedingly rare. Mentors tend to think of themselves as technical advisors, or administrators, or CEOs, or anything other than intellectual models for their students. And their students suffer as a result, because even with the trappings of productive mentorship — meetings, answered emails, edits on papers — they are not learning the only skill that plausibly differentiates PhD-holders from everyone else.

Learning how to ask Goldilocks questions has always been important. But I should note, too, that the advent of generative AI is making it more important than ever. GPT-5 doesn’t possess PhD-level intelligence. Not because it hallucinates or can’t label a map; because it can’t ask a good question to save its nonexistent life.

But it can do statistics, and program, and make figures, and review literature, and do basically everything else that a PhD program ostensibly teaches you how to do. And five years from now, it will do all of those things better than me or any other scientist in history.

So a graduate student in the modern age must cultivate the one ability that has always distinguished scientists from everyone else, and now distinguishes scientists from GPT-5. They must learn how to ask Goldilocks questions.

To all the advisors out there, then — you must teach them. Take your students for a walk.

Thanks for reading Unpublishable Papers! Subscribe for free to receive new posts and support my work.

1

Yes, this is slightly autobiographical.

2

In my experience, the small → big → just right path is more common. One path isn’t better than the other, but most students seem to start out with small questions before broadening their scope of inquiry, rather than the inverse.

3

I want to make it abundantly clear that the below picture is not meant to suggest that my advisor and I are akin to Einstein and Gödel. The point is that going on walks is a good way to come up with interesting questions. Which I do, and so did Gödel . . .

(I take it back, I’m smarter than Gödel, because I both go on walks AND know that most of my food isn’t poisoned).