MBA Thesis  •  Chapter II

Theoretical Framework

This chapter sets out the prior theory on which this thesis stands. The paper addresses a question about people and machines: what happens to human judgment when firms hand over cognitive work to agentic AI? That question is not new. Researchers have studied human-automation pairs for half a century, and organizational scholars have studied the kind of tension this paper diagnoses for just as long. This chapter gathers both traditions and turns them into the lens the later chapters apply.

The chapter has five parts and a synthesis. Part II.1 introduces the foundational human-automation canon and maps each strand to a cluster or to the coordinate spine. Part II.2 sets out the Dynamic Equilibrium Model (DEM) of Smith and Lewis (2011), the full theory of how a dilemma differs from a paradox, and why the tensions in this paper take the shape of paradoxes. Part II.3 introduces jagged intelligence as the substrate that produces those paradoxes. Part II.4 introduces the four-level scale of automation as the coordinate this paper places every tension on, and the well-posed versus ill-posed lens that explains which problems agentic AI can be trusted with. Part II.5 introduces the How Might We question as a design-thinking technique, so that the analysis chapter applies a tool this chapter has already established.

The scope of this paper is one supercluster, Outsourcing Human Cognition, and its five clusters: Trust Miscalibration, Automation Bias, Automation-Induced Complacency, Automation-Induced Skill Atrophy, and Cognitive Debt Accumulation. The theory in this chapter is general, so it would serve a wider study. Here it is read through the single question of outsourcing human cognition, so that the framework motivates the five clusters and nothing more.

II.1 · Foundational canon

The foundational human-automation canon

The vocabulary that lets researchers reason about agentic AI did not begin with large language models. It began half a century earlier. Once early automation made it possible to put a machine in front of a person, engineers began to ask how the pair would behave. Five bodies of work form what this paper calls the foundational human-automation canon. None of these authors wrote about generative AI. Read forward, each one names a pattern that returns, at higher stakes, once a probabilistic agent sits where a deterministic machine used to sit. Together, they give this paper its analytical tools and its longest record of evidence. Each strand also maps cleanly onto a cluster or onto the coordinate spine.

Sheridan and Verplank (1978), mapped to the coordinate spine. Working on remotely operated undersea vehicles, Sheridan and Verplank proposed the levels-of-automation scale: a ten-step ladder running from fully manual control, where the human does everything, to fully autonomous operation, where the machine decides everything and ignores the human. The intermediate steps split decision authority in two ways: what the machine offers as options, and what it does after the human accepts or rejects one. The lasting point is not the ten steps. The principle is that automation is not on or off. Real systems sit on a spectrum, and the position on that spectrum has consequences. This paper carries the scale forward as a four-zone coordinate, described in Part II.4, and uses it to place every tension. Most of the tensions in the corpus appear at the transitions between zones, as authority shifts from human to machine, rather than at any single fixed point.

Bainbridge (1983) mapped to Automation-Induced Skill Atrophy. In a four-page paper, Lisanne Bainbridge named the ironies of automation. The argument is short. Designers automate the easy parts of a task. The human is left with the hard parts, the rare and high-stakes cases that the automation could not handle. But the human's ability to handle those cases depends on routine practice, and the automation has just removed the routine. So the operator loses competence in exactly the part of the job that still needs a human. Bainbridge put it plainly: the more advanced the automation, the more the system depends on the human operator, and the less prepared that operator is to step in. Her 1983 example was a chemical plant whose normal operations were automated but whose emergency response still required a person. That person was asked to do the hardest part of the job, on rare occasions, with none of the routine practice that would have kept them sharp. The pattern now returns in software, in clinical decision support, and in fraud review. It is the substrate of the cluster Automation-Induced Skill Atrophy, whose human-side pole this paper names skill atrophy held against skill maintenance.

Endsley (1995) and Endsley and Kaber (1997) mapped to Automation Bias. Mica Endsley gave the field a model of situation awareness in three levels: noticing the elements in the environment, understanding what they mean, and projecting what they will do next. Endsley showed that automation does not simply replace the operator's action. It changes the operator's awareness of the situation, and the change is rarely benign. In 1997, with David Kaber, she coined the term "out-of-the-loop problem.” When an automated system performs routine work, the operator's situational awareness degrades. When something then goes wrong, the operator can no longer rebuild a picture of the system fast enough to step in. This is the cognitive substrate on which the cluster Automation Bias sits. The operator's attention budget hits a wall the moment the machine's recommendation arrives, so the operator accepts it rather than checking it. The cluster's human-side pole is cognitive passivity held against operator vigilance.

Parasuraman and Riley (1997) and Parasuraman, Sheridan and Wickens (2000), mapped to Automation-Induced Complacency. Raja Parasuraman and his colleagues gave the field two lasting tools. In 1997, Parasuraman and Riley named four failure modes of human-automation interaction: use, which is appropriate engagement; misuse, which is over-reliance on a system that does not warrant it; disuse, which is ignoring a system that would help; and abuse, which is a designer automating a function that should have stayed human. The taxonomy lets researchers diagnose a specific failure rather than gesture at automation problems as one undifferentiated thing. In 2000, Parasuraman, Sheridan and Wickens added a model of where automation can be applied: to gathering information, to analyzing it, to selecting a decision, or to carrying out an action, with the level of automation set separately at each stage. The cluster Automation-Induced Complacency is the misuse failure mode at scale: a reliable machine earns over-reliance, and the operator stops watching it closely. The cluster's pole is operational, with an over-reliance on the processing speed that the machine delivers.

Lee and See (2004) mapped to Trust Miscalibration. John Lee and Katrina See gave the field the canonical treatment of trust calibration: the match between how much an operator trusts a system and how capable that system actually is. Trusting the above capability leads to misuse, where the operator follows the machine when they should override it. Trust below capability produces disuse, where the operator overrides the machine when they should follow it. Appropriate reliance, the state worth aiming for, needs calibrated trust, which in turn needs an accurate mental model of when and why the machine succeeds or fails. This lands directly on the cluster Trust Miscalibration. Operators of LLM-based decision aids miscalibrate because the failure surface of a probabilistic system is jagged in a way that their mental model, built on the smooth failure surface of older deterministic automation, cannot anticipate. The cluster's pole is miscalibrated trust held against informed reliance.

Read together, the five anchors do one thing this paper relies on at every layer. They give the agentic era a vocabulary that pre-dated it. The patterns observed in 2026, productivity paired with cognitive debt, reviewer skills eroding under the automation they are meant to audit, and trust failing on a jagged capability surface, were named and partly measured before any of the underlying technology existed. This paper does not claim that the canon answers the agentic-era questions. It claims that the canon supplies the primitives by which those questions can be classified.

II.2 · Equilibrium model

The Dynamic Equilibrium Model and the dilemma-paradox distinction

The canon supplies the empirical vocabulary, but it does not by itself tell a manager how to handle the tensions it surfaces. For that purpose, this paper turns to paradox theory in organizational scholarship, specifically the Dynamic Equilibrium Model of Smith and Lewis (2011). The model provides the machinery for distinguishing a tension that demands a forced choice from one that demands a sustained balance. This section provides a full account of that model because the later chapters cite it rather than reteach it.

The dilemma-paradox distinction. Smith and Lewis draw a structural line that everyday business language tends to blur. A dilemma is a tension with two mutually exclusive poles. Choosing one means giving up the other, and the right response is a defensible choice between them. A paradox is a tension whose two poles are interdependent and persistent. Choosing one does not remove the other, and the tension returns after every attempt to settle it by picking a side. The distinction is not a matter of words. The two structures call for different responses. A dilemma is managed through choice logic: clarify the criteria, weigh the trade-off, commit, and accept the cost of the road not taken. A paradox is managed through dynamic equilibrium: accept that both poles must be held together, set the system up so that strengthening one pole does not weaken the other, and treat the balance itself as the thing being delivered. Applying choice logic to a paradox is the failure mode that produces an oscillation: each pole is picked in turn; each choice triggers a reaction from the neglected pole; and the organization spends its energy on a question it has misread at the level of structure.

The four DEM categories. Smith and Lewis identify four categories of paradox that recur across organizational scholarship, each named for the activity that creates the tension. Learning paradoxes arise from the demand to both exploit existing knowledge and explore new knowledge simultaneously. This is the classic tension that March (1991) described, generalized to any system whose short-term performance depends on routines that its long-term performance depends on disrupting. Belonging paradoxes arise from the demand for individual identity and collective identity at once: the operator who must trust their own judgment and trust the team's process, the firm that must hold its values and adapt them to context. Organizing paradoxes arise from the demand for both structure and flexibility at once: process discipline versus process improvement, control versus autonomy. Performing paradoxes arise from competing goals or competing stakeholders: short-term profit against long-term value, compliance against the pace of innovation. The categories are not mutually exclusive, and a single tension can sit across two of them. They impose discipline. Naming which category a tension belongs to clarifies which management response will work with the tension's structure rather than against it.

Where the five clusters sit. The five clusters do not all fall under one paradox category. Three are Learning paradoxes, one is a Performing paradox, and one is an Organizing paradox. The three Learning paradoxes share one shape: a tension between using what a machine already does well and preserving the human capacity that long-term performance still depends on. Trust Miscalibration trades an accurate read of machine reliability against the habit of reading it. Automation-Induced Skill Atrophy trades practice against the convenience of letting the machine do the work. Cognitive Debt Accumulation trades a shared theory of the work against the scale of machine-generated output. In each of these three, the short-term gain rests on a routine whose loss the long-term cost depends on. That is the Learning paradox in its agentic-era form. Automation Bias is a Performing paradox. Its poles are competing goals at the moment of decision: the speed of accepting the machine's output against the accuracy that an independent check protects. Automation-Induced Complacency is an Organizing paradox. Its poles are competing process designs: sustained oversight against the operational reliance that makes oversight feel wasteful. The thread across all five is that each tension pulls one scarce human faculty in two directions at once.

Why are these tensions paradoxes and not dilemmas? The paper's headline finding is that all five clusters are paradoxes. That finding needs a structural explanation, not just a count of cases, and the Dynamic Equilibrium Model provides the framework for it. A dilemma needs one pole that is defensible on its own. A paradox arises when neither pole is defensible on its own, and the two poles are genuinely independent forces. For each of the five clusters, neither pole survives alone. Trusting the machine without reading it lets bad outputs into real decisions. Refusing to trust it forfeits a capability the firm is paying for. Watching every output returns the firm to the speed it was trying to escape. Watching none of them lets silent failures through. Because neither pole is safe to settle into, the tension is a structural paradox, and the management response is to hold both poles rather than choose one. The next section explains what property of agentic AI makes both poles indefensible at once.

II.3 · Jagged intelligence

Jagged intelligence as the paradox-producing substrate

The Dynamic Equilibrium Model identifies paradox as a recurring organizational form, but Smith and Lewis do not identify a substrate that produces it. This paper proposes one for the agentic era: jagged intelligence. This is the deepest answer to why the five clusters come back as paradoxes rather than dilemmas, and it is the argument behind the title.

Karpathy (2024) coined the term "jagged intelligence” to describe a property of frontier AI: a non-smooth capability surface. The same system is superhuman at some tasks and worse than a child at adjacent, trivial ones, and a person cannot tell in advance which task will fall on which side. Models win gold at the International Mathematical Olympiad and score above expert validators on a PhD-level science benchmark. The same models read an analog clock correctly only about half the time, insist that 9.9 is smaller than 9.11, and miscount the letters in a short word. Capability and failure sit side by side, with no smooth gradient between them and no reliable signal as to which is which.

This jaggedness is what turns each tension into a paradox by structural necessity. On a smooth capability surface, a manager could learn where the machine fails, commit to trusting it inside the safe region, and treat the tension as a dilemma with one defensible pole. A jagged surface removes that option. Trusting the machine fails because the gaps bite without warning. Refusing to trust the machine also fails, because forfeiting the region where it is genuinely competent is the larger loss. Neither pole is defensible alone, so the tension cannot be a dilemma, and paradox follows from the substrate itself rather than from operator psychology or organizational politics. This is the structural answer to the question of why this paper finds paradoxes rather than dilemmas: jagged intelligence is a structural property of modern AI, and paradox is the shape a tension takes when a person cannot predict where the substrate will fail.

The same argument explains the one exception in the corpus. Where the jagged gaps are bounded, where the capability surface is narrow, the output is verifiable, and the failures are predictable, the tension trends back toward a dilemma with one defensible pole. This is consistent with the evidence base, in which the single dilemma-typed item falls under Automation Bias in a setting where verification is tractable. The substrate argument is therefore falsifiable rather than decorative: it predicts where dilemmas should reappear, and the corpus matches the prediction.

This is the substrate extension to Smith and Lewis. They describe paradox as a recurring shape but name no cause. This paper proposes jagged intelligence as the agentic-era cause that generates the paradoxes the model describes. The tensions are paradoxes by a system property, not by culture.

One citation distinction matters here because two similar terms describe different things. Karpathy's jagged intelligence (2024) refers to the internal capability distribution of the model itself, that is, what the model can and cannot do. Dell'Acqua and colleagues (2023) describe a jagged frontier, which is the external boundary of worker productivity, that is, which tasks an AI-assisted worker performs better or worse on. The two are related but not the same. This paper uses Karpathy's jagged intelligence as the substrate-level cause of the paradoxes, and reserves Dell'Acqua's jagged frontier for the worker-productivity boundary. Jagged intelligence itself belongs to the AI side of the wider thesis rather than to this paper's five human-cost clusters. It is referenced here as the upstream cause explaining why the five clusters are paradoxes, rather than being added as a sixth cluster.

II.4 · Coordinate spine

The coordinate spine and the well-posed lens

Two further instruments place each tension in space and indicate which problems agentic AI can be trusted with.

The four-level scale of automation. Sheridan and Verplank's ten-step scale, introduced in Part II.1, is carried over into this paper as a four-zone coordinate system: Human, AI-assisted, Human-in-the-loop, and Autonomous. The four zones preserve the diagnostic value of the original while remaining legible across the entire corpus. Each tension is placed on this coordinate, and the placement matters because most tensions surface at the transitions between zones rather than inside any single zone. The cost appears as authority shifts from human to machine, or as the human-in-the-loop role becomes nominal, a sign-off in name only. The coordinate gives those transitions a vocabulary, so the analysis can say where on the human-to-autonomous path a given tension bites.

The well-posed versus ill-posed lens. Ming offers a second lens that complements the canon: the distinction between a well-posed problem and an ill-posed one. A well-posed problem has a clear specification, a verifiable answer, and a stable definition of success, so a capable system can be trusted to handle it, and a human can verify the result. An ill-posed problem has no clean specification, no single checkable answer, and a definition of success that shifts with context, so handing it to a machine produces output that looks finished but cannot be verified at a glance. The lens pairs with jagged intelligence. A jagged system is most dangerous for ill-posed problems, where humans have no easy way to distinguish a competent answer from a confidently wrong one. It is safest on well-posed problems, where the answer is checkable and the jagged gaps, if they appear, are caught. The five clusters in this paper all involve outsourcing cognition on problems that are at least partly ill-posed, which is why each one resists a clean forced choice.

II.5 · How Might We

The How Might We question as a design-thinking technique

The final instrument this chapter installs is the How Might We question, so that the analysis chapter applies a tool already established here. A How Might We question is a design-thinking technique used in product and service design to open a problem rather than close it. It takes a difficult situation and reframes it as an invitation: not which option to choose, but how we might achieve a goal given the constraints. The phrasing is deliberate. How signals that a workable approach exists and has not yet been found. Might keeps the space open rather than committing to one answer. We frame the work as shared.

The technique fits a paradox for a precise reason. A paradox must be managed, not solved, because neither pole can be safely chosen. A How Might We question presupposes that both poles coexist and asks how to manage their interaction, which is exactly the response the Dynamic Equilibrium Model prescribes for a paradox. It shifts the question from which pole we pick to how we might hold both poles so they reinforce each other rather than erode each other. For the five clusters in this paper, the recurring shape of such a question sets up a cross-check: how might we design the work so that the machine's strengths check the human's weaknesses, while the human's strengths check the machine's. Because the question presumes coexistence rather than choice, it turns a binary decision into a design space a manager can act in.

Introducing the technique here matters for what follows. When the analysis chapter ends each cluster with a How Might We question, it is not producing a loose end or a template flourish. It is applying a design-thinking move this chapter has already established, one that treats the paradox as a starting point for design rather than a dead end. The paper diagnoses the paradox and then hands the reader an open door to act through.

II.6 · Synthesis

Synthesis

This chapter has assembled the borrowed theory on which the paper stands. The foundational canon supplies the diagnostic primitives and maps each strand to a cluster or the coordinate spine: Sheridan and Verplank on the coordinate spine; Bainbridge on skill atrophy; Endsley and Kaber on automation bias; Parasuraman and colleagues on complacency; and Lee and See on trust miscalibration. Cognitive Debt Accumulation is the agentic-era cluster with no anchor in this pre-LLM canon. The Dynamic Equilibrium Model supplies the structural distinction between a dilemma and a paradox, the four categories of paradox, and the basis for typing each cluster as a Learning, Performing, or Organizing paradox. Jagged intelligence supplies the substrate that makes both poles indefensible at once, which is why the tensions are paradoxes by structure. The four-level scale of automation and the well-posed lens place each tension in space and say which problems a machine can be trusted with. The How Might We question provides the design-thinking move that turns each paradox into a space to act.

The next chapter operationalizes these instruments. It develops the Dichotomy Probe, which routes each tension to a verdict, and the Equilibrium Mapping, which renders a paradox as a balance of poles.

Bibliography

References

  • Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775-779.
  • Dell'Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier (Harvard Business School Working Paper No. 24-013).
  • Endsley, M. R. (1995). Toward a theory of situation awareness in dynamic systems. Human Factors, 37(1), 32-64.
  • Endsley, M. R., & Kaber, D. B. (1997). Out-of-the-loop performance problems and the use of intermediate levels of automation. Human Factors, 39(2), 230-253.
  • Karpathy, A. (2024). Jagged intelligence.
  • Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80.
  • March, J. G. (1991). Exploration and exploitation in organizational learning. Organization Science, 2(1), 71-87.
  • Ming, V. (Robot-Proof, year pending verification).
  • Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253.
  • Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics, Part A, 30(3), 286-297.
  • Sheridan, T. B., & Verplank, W. L. (1978). Human and computer control of undersea teleoperators (Technical report). MIT Man-Machine Systems Laboratory.
  • Smith, W. K., & Lewis, M. W. (2011). Toward a theory of paradox: A dynamic equilibrium model of organizing. Academy of Management Review, 36(2), 381-403.