What Makes a Good Question in the Age of AI

In a recent TED talk titled One Thing to Teach in the Age of AI, innovation practitioner Bryan Cassady shared a story about an exam created by one of his colleagues. The exam question was deceptively simple: “Prove you’ve learned the material in this course by giving me five good questions.” Many students struggled with that, and certainly nobody received a perfect score.

The question wasn’t hard because of what it asked for; it was hard because of what it didn’t ask for: answers.

With the wide availability of modern generative AI models, which Dario Amodei describes as “a country of geniuses in a data centre” and Noah Smith likens to be the next evolution of a centralised all-knowing world mind, many different forms of knowledge are now virtually free, unlimited, and accessible on demand to anyone and everyone. From the perspective of a university lecturer, we have certainly known for a while that having well-known answers to standard questions is no longer a good gauge of a student’s understanding in this brave new world of generative AI. I am intrigued by Cassady’s argument that being able to ask good questions is now the next-best test of understanding.

But what makes a good question in the age of generative AI? 

Here’s a mental model that may be useful.

In this diagram, I first separate out the Knowable Facts and the Unknowable Statements. (The latter is a hat tip to Godel’s Incompleteness Theorem and can also be understood more simply as NP-hard-type problems where the effort required to establish the definitive correctness of a statement is beyond the finite time measurable in human lifespans.) Within the Knowable Facts circle, I separate out basic, atomic facts and more complex statements that require inference. For clarity, I also used a purple line to partition the circle into True Statements and False Statements. In the above, one should interpret “statements” as terms in higher-order logic, which means the statements can be probabilistic and inference includes both logical inference (e.g. modus ponens, universal rule, etc) as well as probabilistic inference (e.g. Bayes rule, marginalisation, etc.)

On top of that basic map I overlay the current but constantly evolving human knowledge — what humanity collectively knows today — and the current and also constantly evolving AI jagged frontier.

I think we can distinguish between three classes of questions that are very loosely tied to the Bloom hierarchy. 

1. Easy Questions

  • Characteristics: These are simple lookup queries on core, settled facts.
  • Human Effort: Asking such questions requires very little effort from people, and they certainly do not need critical thinking or domain context.
  • Efficiency: Using AI to answer such questions is effectively a waste of compute. These easy questions can be answered more reliably and with less resources using standard search engines and databases (e.g. Wikipedia).

2. Reasonable Questions

  • Characteristics: These are questions whose answers have to be inferred using non-trivial logical reasoning steps. 
  • Human Effort: Asking such questions requires some level of curiosity and domain knowledge from people, and understanding the answers that come back from AI requires some level of sophistication to comprehend, for example, domain-specific reasoning methodologies and/or mathematical proofs.
  • Efficiency: Using AI to answer such questions can help non-experts get up to speed in a domain quickly. However, domain specialists can usually get answers to these questions faster using specialised tools. 

3. Good and Deep Questions

  • Characteristics: These are questions where the answers sit in the region where human domain expertise meets the AI jagged frontier. Good questions are the ones that are just beyond current human knowledge but can be answered with the best AI models today. Deep questions are good questions that lie right at the border between Knowable Facts and Unknowable Statements. 
  • Human Effort: Asking such questions requires strong domain knowledge to frame high-value hypotheses and strong critical judgment to distinguish genuine insight from plausible noise.
  • Efficiency: People who ask good and deep questions use AI as an intellectual sparring partner to explore edge cases, stress-test logic, and advance the state of human knowledge. These are the most beneficial way of using AI. 

In my view, a student who can ask good and deep questions demonstrate beyond doubt their mastery of a subject matter. However, this is likely too high a bar for undergraduates. But I think we should strive to get as close as practical to this limit in the courses we teach in universities. 

As Cassady and others have argued, the premium in the age of AI has shifted entirely from knowledge retention to question formulation and answer evaluation. The winners won’t be those with the fullest heads, but those with the sharpest questions and the judgment to navigate the jagged human-AI knowledge frontier. We have to start inculcating such skills in students as early as possible.


Leave a comment