
Thought Leadership
September 9, 2026
Preview/Summary: In the world of education, what gets assessed gets taught. If our credentialing systems continue to test elements of the job that AI can now perform, educators will continue to prepare candidates for an obsolete workplace.
In more than 40 countries, judges have now had to rule on the same strange failure: a lawyer filed a document citing cases that do not exist. A public database of those decisions has logged more than 1,900 rulings globally.
Almost none of the lawyers were incompetent by the standard their qualification measured. They knew the law. What nobody tested them on was whether they could recognize a fluent technology being confidently wrong. That gap in what assessments measure and the skills needed to be competent in AI-assisted work will increasingly affect every profession that licenses people to be trusted.
The promise of high-stakes professional testing is straightforward: It should ensure that future practitioners have acquired the skills they need to serve the public competently, from accounting to law to healthcare.
But the integration of AI in the workplace changes what it means to be a competent professional in fundamental ways. How much can we trust a practitioner who lacks reliable judgement and a strong grasp of how to collaborate with AI in today's professional environment?
These developments have huge implications for providers of professional assessments, including Kaplan. Our work suggests professional credentialing pathways will need to adapt in significant ways to measure how candidates harness AI to create better outcomes. We must move beyond measuring baseline content knowledge and instead focus on how well candidates deploy AI and critically evaluate its output.
To do this effectively, we first need to distinguish between two different cognitive behaviors that are often linked to the use of AI tools. The first, often known as cognitive offload, refers to a valuable and legitimate everyday skill that should be encouraged: the efficient outsourcing of routine tasks. Think of using a calculator to speed up basic computation. Offloading chores to AI is very similar. It allows a professional to free up mental capacity for higher-order strategy that requires nuance and critical thinking.
By contrast, the second behavior, cognitive laziness, is much more problematic. It’s what happens when a user completely disengages their critical faculties and follows AI instructions unquestioningly. Abdicating decision-making to technology this way is a bit like a driver who follows GPS instructions to the letter without looking to see whether they are being steered down a dead-end street.
AI is changing how people learn, and whether they remember learning it at all. A striking four-month study from the MIT Media Lab found that when individuals relied heavily on generative AI, the brain rhythms associated with deep memory formation actively declined. By the final session, 78% of participants who relied on ChatGPT could not quote anything from material they had produced just minutes earlier.
Cognitive laziness directly relates to the much-discussed crisis of deskilling among junior professionals. As AI automates so many entry-level tasks, from compiling draft financial reports to writing routine code, junior workers are losing the traditional training grounds where they build foundational skills and professional judgment. Denied the chance to develop the practical, “human-add” abilities that will make them tomorrow’s experts, they’re less likely to challenge the machine and more likely to succumb to cognitive laziness.
The good news is that assessment can play a useful role in defending against this hazard. The path forward is to teach, then evaluate, the practical cognitive skills required to supervise a machine and retain the human comparative advantages. These capabilities include:
Critical evaluation of AI outputs. The ability to spot subtle fabrications, outright hallucinations, and basic factual or contextual errors.
Justified deviation. The confidence and independent judgment to challenge an automated AI recommendation and to defend one’s human, professional decision.
Prompt calibration. The capacity to give directions and refine generative AI outputs instead of passively accepting the first draft.
Metacognitive awareness. The ability to recognize when to rely on fast, automated tools and when to apply human intervention.
Next comes an even harder task: designing the assessments themselves. That work is mostly ahead of us, and it has to satisfy multiple objectives.
First, certain baseline knowledge needs to be measured without any access to AI. Next, testing a candidate's ability to use AI will of course require that AI tools be available in the testing environment. In medical certification, that could mean scoring how a candidate handles an AI-generated differential diagnosis — or whether they catch what a diagnostic algorithm missed.
Then there's the need to stay grounded in specific professional disciplines — because mastery of foundational domain knowledge is crucial to identifying whether AI is doing a good job. A legal professional can’t detect a flawed AI-generated contract clause, nor can an accountant catch a defect in an automated financial audit, without deep, field-specific expertise.
Many open questions remain, including whether and how to train AI avatars to behave like seasoned human examiners, and how to ensure that tests remain sufficiently standardized that their results are reliable and comparable.
Even a redesigned exam solves only half the problem, because it still assumes that competence measured once stays measured.
Most regulated sectors still operate on a dual credentialing model: a rigorous, front-loaded high-stakes assessment to gain an initial license to practice, followed by lower-stakes, often unverified, requirements for continuing professional development. From medicine to law and beyond, this approach assumes that initial competence remains permanent and that passive exposure to information guarantees ongoing fitness to practice. In a fast-moving world where professional demands shift in months rather than decades, that model no longer provides reliable public assurance.
There must be an increasing emphasis on ongoing competence, regular recertification and lifelong professional development. To maintain genuine public trust, recertification must be evidenced via valid, reliable assessment events throughout a professional's career — not by a digital badge awarded for attending a webinar, which remains relatively meaningless. The question is no longer whether we assess at a single point in time, but how we structure a robust architecture of lifelong validation.
This transition in how assessments are designed, and what they're intended to measure, will take time and effort to navigate thoughtfully. But the direction is clear. The role of the high-stakes exam is shifting, from a one-time gateway at the start of a career to a recurring anchor that validates professional judgment in a faster-moving world.
Zoe Robinson is Managing Director, Kaplan Assessments.
Should candidates be allowed to use AI in professional exams?
It depends on what the stage is measuring. Baseline knowledge must be tested with AI unavailable, or the exam measures the tool rather than the candidate. Applied judgment must be tested with AI available, or the exam measures nothing that resembles modern practice.
What AI skills should employers and licensing bodies test for?
Key skills include critical evaluation of AI output, justified deviation from AI recommendations, prompt calibration, and metacognitive awareness of when to delegate. Each needs to be defined against an observable behavior to be scorable.
What is the difference between cognitive offload and cognitive laziness?
Cognitive offload is delegating routine work to free capacity for judgment — a skill worth encouraging. Cognitive laziness is delegating the judgment itself, accepting output the person is not equipped to evaluate.
Does the EU AI Act apply to exam and assessment systems?
Yes. AI systems used to evaluate learning outcomes or administer exams fall under Annex III as high risk. The AI Digital Omnibus, in force 24 July 2026, moved most stand-alone Annex III obligations to 2 December 2027.
Not finding what you’re looking for?