Beyond the Sniff Test: AI, Human Judgement and the Future of Expertise

A few months ago, I wrote about applying the Sniff Test to AI-generated content.

The premise was simple. As Generative AI becomes increasingly capable, we should be diligent in maintaining a degree of healthy scepticism:

  • Does this align with what I already know?
  • Can I explain it in my own words?
  • Is there a logical thread?
  • If this came from a brilliant but junior analyst, would I still double-check it?

The article resonated because most experienced professionals have encountered the same phenomenon. AI can produce highly persuasive answers, confidently presented and logically structured, while still containing flawed assumptions.

The ability to recognise those issues often comes down to experience. Which raises a much bigger question.

“If experience is required to apply the Sniff Test,

where will that experience come from in an AI-enabled future?”

The Human in the Loop Assumption

Many AI governance frameworks contain a common recommendation:

“Keep a human in the loop.”

It sounds sensible to allow AI to generate reports, recommendations, code, analysis, designs and content, but require a human to review the output before it is used.

The assumption is that human oversight reduces risk, but this assumption relies on

  • The human has enough experience to provide meaningful oversight.
  • The human is not so fatigued in providing oversight that they are simply “box-ticking”
  • The human can understand the “rules” followed to create the output.

Why Experience Matters

Most professionals do not develop judgement through being taught. They developed it through repetition as junior analysts, graduate developers, junior consultants gathering requirements or helpdesk agents resolving issues. Over time, they begin to recognise patterns, and they learn what right and wrong look like.

Eventually, they acquire the ability to review the work of others and provide guidance that people actively seek out, not because they are smarter than everyone else but because they have accumulated enough experience.

The Disappearing “Experience” Apprenticeship

This is where AI presents a unique challenge. Previous waves of technology largely automated physical effort; AI is beginning to automate cognitive effort.

The activities being replaced are often viewed as low-value work, yet they are incredibly valuable as they are the apprenticeship through which future subject matter experts are created.

The irony is that the more successful AI becomes at performing entry-level cognitive work, the fewer opportunities there may be for people to develop the experience required to supervise it.

Organisations could find themselves increasingly dependent on AI at precisely the same time that experienced reviewers become harder to find.

The Rise of the Reviewer

The future junior role will be very different; future professionals will spend less time generating information and more time validating it by testing, refining and governing AI-generated solutions.

Building Better Guardrails

The AI Reviewer will have several options, some of which include:

  • Using multiple AI models to independently solve the same problem and compare results.
  • Benchmarking outputs against known datasets and historical outcomes.
  • Testing recommendations against predefined business rules.
  • Maintaining audit trails that explain how conclusions were reached.
  • Continuously measuring the accuracy and effectiveness of AI-generated outputs over time.
  • Using independent human and AI review processes to challenge recommendations before decisions are made.
  • Applying a Sniff Test by requiring recommendations to be explainable, evidence-based and consistent with known business outcomes.

The Black Box Problem

The problem is that many guard rails assume that the output is deterministic; however, AI introduces another challenge: most Large Language Models are effectively black boxes. For the average business user, understanding the internal workings of a billion-parameter model is unrealistic.

This creates an important shift in thinking: perhaps the goal should not be to completely prevent AI from making mistakes, which may be impossible, given that AI systems are probabilistic.

The challenge therefore becomes less about eliminating every error and more about detecting, validating and correcting errors before they create harm. The objective is not to understand every internal decision made by the model; The objective is to recognise when the outcome should be questioned.

Beyond the Sniff Test

The more I use Generative AI, the clearer I see the challenge to be whether organisations can redesign learning pathways, governance frameworks and validation processes quickly enough to ensure AI is successfully implemented and governed.

The organisations that succeed will not necessarily be those that deploy the most AI. They will be those who become best at developing judgement, validating outcomes and creating the next generation of AI Reviewers.

The organisations that adapt successfully will discover that the most valuable capability is no longer content creation, but judgement.

Scroll to Top