AI versus the MBA
What skills will entry-level knowledge work require and how should MBAs and undergrads be trained?
Imagine you’re a first-year McKinsey analyst or associate. On Monday morning, your manager says. “Our client is a telecom company that wants to expand into software. Develop a thesis for which adjacent software markets to consider and identify acquisition targets in each space. Come back with a specific market and acquisition recommendation by Friday.”
You are neither told which framework to use nor is there a formula to apply. Instead, you gather data, interview people, analyze competitors, build financial models, estimate market sizes, and eventually assemble a slide deck that argues for a course of action. Then your manager tears it apart. You revise. The partner tears it apart again. Over time, you learn not just how to analyze a business problem, but how to think like an experienced C-suite decision-maker.
Now replace McKinsey with Amazon, Goldman Sachs, Blackstone, General Mills, or any Fortune 500 company. The details differ based on whether you are in strategy or operations or marketing or another function, but the pattern is the same. Much of the early career of knowledge workers in business consists of taking an ambiguous problem, applying analytical tools, and producing a recommendation that someone more senior can evaluate, sharpen, and implement.
In new research, my coauthors and I raise a question: what happens when AI becomes remarkably good at the entry-level tasks that most college and MBA graduates perform at the workplace. And what does it mean for the future of work and business education?
We asked AI to do an MBA’s job
Most AI benchmarks tell us whether a model can recall facts, solve math problems, write code, or answer exam questions. Those are useful capabilities, but they are a poor reflection of what most white-collar professionals actually do.
Knowledge work is messy. Managers rarely receive neatly packaged questions with objectively correct answers. They receive incomplete information, conflicting objectives, and uncertain forecasts, and have to juggle organizational politics and pressure to make decisions fast.
In business schools, the case method has become the signature tool for training students for exactly this kind of knowledge work. Rather than teaching theory alone, instructors place students inside business situations (some real and others fictional, but all realistic). Students have to wrestle with the same uncertainty managers face when making decisions: information is incomplete, stakeholders disagree, and every option carries trade-offs. There is usually no single correct answer.
Together with my coauthors, we built BusinessCaseBench, a benchmark based on 238 business school cases spanning 18 disciplines, from finance and accounting to strategy, leadership, operations, marketing, entrepreneurship, ethics, and international business. The benchmark contains 615 open-ended questions, each graded against expert-written solutions using detailed scoring rubrics. Access to these cases is gated and they are unlikely to be in LLM training datasets. In particular, the sample solutions are only accessible to verified instructors and even more unlikely to be in LLM training data
The results surprised us. Today’s leading AI models score above 87% against instructor rubrics. Within one model family (OpenAI), performance improved by roughly 23.3 percentage points in just two years. AI is no longer merely answering factual finance questions or calculating discounted cash flows. It is synthesizing complex situations, weighing alternatives, and producing strong business recommendations across almost every business discipline we examined.
But there is one nuance.
We evaluated responses in two ways. The first awards partial credit based on the fraction of rubric criteria satisfied. The second asked a stricter question: did the model produce a complete answer, satisfying every important criterion identified by the instructor? Here, performance dropped to under 50%. AI often generated strong overall analyses while overlooking an important stakeholder, omitting a key assumption, failing to discuss an implementation challenge, or stopping just short of a truly decision-ready recommendation.
Fig 1. In standard scoring, scores for OpenAI family of models went up by 23.3%
In short, AI is increasingly good at producing exceptional first drafts. It is not yet consistently good at producing gold-standard final drafts i.e. matching instructor solutions along all dimensions (To be clear, I doubt MBAs at top business schools will score much higher, on average, against this stricter criterion). And that has big implications as I explore next.
If AI can do the cases, what is the MBA for?
For decades, much of business education has focused on teaching students to generate sound analyses to support decision-making under uncertainty. Increasingly, AI will do that in seconds. The scarce skill now is verifying whether a draft — completed by someone else — is actually good enough to support an important decision. That is a harder skill. And one that has historically been developed by doing the work oneself over and over. As my colleague Michael Roberts put it in our email exchange, “[it] used to require years of building DCF models to gain sufficient experience so one could pick up someone else’s model and assess it.” If AI systems perform more of the production work, firms may need new apprenticeship models that develop verification and model supervision earlier in career paths.
The other question is where will value shift at the workplace. BusinessCaseBench measures a central part of knowledge work: the production of structured analysis, but it does so under relatively clean, single-turn conditions. In organizations, the same decision problem often comes with differing opinions, organizational politics, and implementation consequences that more experienced managers have to wrestle. If AI produces strong first drafts, I believe value will likely migrate upstream toward the selection of which question is worth answering, and downstream toward building consensus, adapting recommendations to organizational context, and owning the consequences of decisions. Expecting entry-level workers to be able to do this is a tall ask but that may in fact be what they have to do to add value over and above what AI can do.
The educational problem for professors like me shifts from whether students can generate acceptable analysis to whether they can evaluate, stress-test, and improve one created by someone else. This shift does not make foundational analytical skills obsolete. Effective verification requires domain knowledge to identify faulty assumptions, omitted stakeholders, weak causal claims, numerical errors, and incomplete recommendations. The instructional burden for professors in fact increases: students must learn both how to perform the work and how to evaluate AI-generated work. Business schools need to add repeated AI-in-the-loop exercises in which students audit flawed or incomplete analyses, compare competing recommendations, and recognize what a strong answer looks like.
There are some interesting implications for model training as well. Many of today’s most famous benchmarks are approaching saturation. Measuring future progress will increasingly require benchmarks grounded in economically meaningful work rather than narrow factual recall or math reasoning. AI labs have been spending big $$$ buying data on typical work performed by consultants, bankers, marketers, etc through post-training data providers such as Turing, Mercor, and AfterQuery. Clearly, Fig 1 shows this has paid off. But we do find a nice subset of open-ended analytical reasoning problems that remains challenging for frontier models. These harder cases provide a natural target for future model development.
Fig 2. Across disciplines, there remain questions that are hard for frontier models (especially so under complete answer scoring but also under partial scoring)
If you are an experienced exec, let me know how some of the above findings and general trends are impacting hiring at your firm. If you are a recent graduate, share your thoughts on what kinds of questions and concerns this raises for you.
Links to the paper and the project website.
In other news, How do you market to an AI customer is a recent HBR article of mine in which I explore how AI is transforming the customer buying journey and why it would be a mistake for marketers to assume that AI is just another channel through which your customers are finding you. Instead, the new problem for marketers is how to inform and influence AI agents as they do research or make decisions on behalf of your customers. My recent keynote at the just-concluded CommerceNext conference also explored the same theme. More updates on that research in a future newsletter.




The gap between standard scoring and complete answer scoring for Gemini Flash is striking. What do you think explains that model losing so many points on completeness specifically?
MR.KARTIK HOSNAGAR, John C. Hower Professor, Renowned Author, Professor of Operations, Information and Decisions,
Co-Director, Wharton Human-AI Research,Professor of Marketing,in Creative Intelligence post—AI versus the MBA—navigates the upcoming AI economy and shares deep research based insights on—What skills will entry-level knowledge work require and how should MBAs and undergrads be trained?
The post is worthy of a complete book and contains latest Research and insights by MR.KARTIK on the intersection of AI and business and draws on:
# MR.KARTIK HOSNAGAR’s new research,with his coauthors that raises a question: what happens when AI becomes remarkably good at the entry-level tasks that most college and MBA graduates perform…
## MR.KARTIK HOSNAGAR’s BusinessCaseBench, a benchmark based on 238 business school cases spanning 18 disciplines,...
### MR.KARTIK HOSNAGAR’s recent HBR article—How do you market to an AI customer…
####MR.KARTIK HOSNAGAR’s recent keynote at the just-concluded CommerceNext conference...