MBA Thesis  •  Chapters V–VI

Discussion and Conclusion

This chapter steps back from the per-cluster analysis and asks what the result means. Chapter IV returned a verdict for each of the five clusters that make up Outsourcing Human Cognition: Trust Miscalibration, Automation Bias, Automation-Induced Complacency, Automation-Induced Skill Atrophy, and Cognitive Debt Accumulation. All five classified as paradoxes. This chapter draws out what follows from that. Section 5.1 sets out the derivatives, the implications that fall straight out of paradox-dominance for the corpus. Section 5.2 turns those into four concrete recommendations for managers. Section 5.3 places the findings back inside the theory introduced in Chapter II. Section 5.4 states the limits of the work and three ways the central claim could be proved wrong.

V · Discussion · 5.1

Derivatives

The result is one number, and it is best stated with care. Among the documented capability-cost tensions in the corpus, 61 of 62 instances are classified as paradoxes, which is 98.4 percent. At the cluster level, the result is cleaner still: all five clusters classify as paradoxes, and none as dilemmas. The single dilemma-type evidence falls under Automation Bias. That distribution is the first thing the instrument produced for this corpus. It is not the thesis's durable contribution, which is the instrument itself. The lone dilemma matters too, because it shows that the Dichotomy Probe still discriminates rather than waving every tension through as a paradox.

The finding should be read in two ways at once. The first reading is direct. When firms outsource human cognition to agentic AI, the recurring tensions they are asked to weigh are overwhelmingly persistent. They do not clear when a decision is made. They push back from the suppressed side and re-emerge in the next quarter's numbers. A business leader who reads only the productivity figure and decides that the AI tool is working has not taken a wrong measurement. They have taken it on the wrong timescale. Productivity and cognitive engagement move in the same direction in the short run and in opposite directions over weeks (Cui et al., 2025; Kosmyna et al., 2025). The paradox-dominance figure measures, in aggregate, how often this timescale collision is the real shape of the tension rather than surface noise.

The timescale collision is sharper here than the headline conveys, because the five clusters expand along one axis. Trust Miscalibration shows up per task, when an operator decides whether to rely on a given output. The cost then runs through attention at the moment of decision (Automation Bias), through sustained vigilance across a shift (Complacency), through practice across a career (Skill Atrophy), and finally through the shared understanding a team holds about its own work over a product lifecycle (Cognitive Debt). Each cluster is a consequence of the one before it. The compressed version is that the same margin of reliability that makes the system worth deploying is what makes verification, vigilance, and skill all decline. A short-run productivity gain and a long-run cognitive cost are not two separate stories. They are the early and late readings of one mechanism.

The second reading is harder to ground because the necessary comparison is missing. The foundational human-automation canon, drawing on Sheridan et al. (1978), Bainbridge (1983), Parasuraman and Riley (1997), Endsley (1995), and Lee and See (2004), surfaced the same mechanism families that this paper names: the irony of automation, automation-induced complacency, trust miscalibration, and the out-of-the-loop problem. The canon did not classify those tensions as dilemmas versus paradoxes in the Smith and Lewis (2011) sense. It more often described a single mechanism in tension language without forcing a verdict on whether the tension could be settled by choice or only by co-management. So the comparison is partial. We cannot say with confidence that the prior automation era was dilemma-dominant. What we can say is that the canon's substantive findings, retested against the Dichotomy Probe in this corpus, overwhelmingly classify as paradoxes, and the underlying mechanisms are recurring rather than new. The agentic era did not invent these five tensions. It scaled them, accelerated them, and made them simultaneous across more decision-makers at once.

Two implications follow. The first concerns the kind of decision-support tooling firms will need. Forced-choice frameworks handle dilemmas well. Build-or-buy, centralize-or-federate, and deploy-or-hold all assume that picking a side settles the matter. They handle paradoxes badly because the underlying tension does not resolve when a side is picked. The result suggests firms need a different kind of framework: one for holding two real goods in productive tension over time, designed for sustained balance rather than a final choice.

The second implication concerns the lone dilemma and why it sits where it does. The single dilemma-typed evidence in the corpus falls inside Automation Bias, in a case where verification is tractable: the operator can check the output, the failure is visible, and the gap in the system's competence is bounded. This is consistent with the structural account of why these tensions take paradox shape at all. Frontier AI has a jagged capability surface (Karpathy, 2024): it is superhuman on some tasks and worse than a child on adjacent trivial ones, and which is which cannot be predicted in advance. A dilemma needs one pole that is defensible on its own. A paradox arises when neither pole is defensible alone. On a jagged substrate, trusting the system fails because the gaps bite, and refusing to trust it fails because forfeiting the competent regime is the larger loss. Neither pole survives, so paradox follows by structural necessity. Where the jagged gap is bounded by a narrow surface and a verifiable output, the tension tends back toward a dilemma. The lone Automation Bias dilemma sits in such a bounded pocket. This dilemma-collapse condition is testable and falsifiable, and the analysis presents it as a theoretical contribution rather than an incidental observation. It predicts that the tension drifts from paradox toward dilemma as its verification surface narrows.

V · Discussion · 5.2

Recommendations

The finding points business leaders toward a small set of operational shifts. These are not a checklist. They are reorientations that change which question gets asked first when a new tension around outsourcing cognition surfaces inside the company.

Diagnose before deciding.

Before a leadership team commits to an automate-or-not question, the prior question is whether the tension is a dilemma or a paradox. A dilemma is a case where the evidence leans one way rather than the other, so one pole is defensible. A paradox is where both poles are real and load-bearing, and neither reduces to a single choice. The Dichotomy Probe is the diagnostic, and its seven steps are documented in full so that any team, not only researchers, can run it on a candidate tension and reach a defensible verdict. In this corpus, the probe returns a paradox for all five clusters. That result is itself decision-relevant. It tells the firm that further debate over which pole to pick is a waste of energy, and that the productive question comes after the verdict, not before.

Design for both sides, not for one.

Once a tension is classified as a paradox, the next question is not which pole to choose. It is about designing the work so that the two poles check each other rather than erode each other. This is the role of the How Might We question, a design-thinking move that presupposes both poles coexist and asks how to manage their interaction.

Because a paradox must be managed and not solved, the How Might We is not a loose end. It is an open door, an invitation for the business leader to act within a design space rather than settle for a forced choice. Each cluster carries its own How Might We, drafted in Chapter IV, and the shape is shared across the five.

  • For Trust Miscalibration: how might we surface model behavior in deployment so that operator trust tracks system reliability rather than drifting away from it?
  • For Automation Bias: how might we surface the AI's confidence so that it triggers a real verification step rather than quiet acceptance?
  • For Complacency: how might we structure oversight so that operator attention is required rather than treated as a mere ceremony?
  • For Skill Atrophy: how might we organize AI-augmented work so that core skills stay in active use rather than decaying through disuse?
  • For Cognitive Debt: how might we sequence assisted and unassisted work so that shared understanding is rebuilt across sessions rather than accumulated as a deficit?

Each question moves the business leader from which side wins to how the two sides hold each other in tension over time.

Watch the architecture, not the operator.

This recommendation most directly contradicts what firms usually do. When a firm meets these cognitive-passivity tensions, the standard response is a training program aimed at operator discipline: stay alert when using AI tools, verify outputs before accepting them, do not let the AI think for you.

The corpus evidence is consistent that these programs do not survive contact with the underlying dynamics. The Klingbeil et al. (2024) experiment showed that subjects given an AI advisor performed worse than unassisted controls and overrode their own contextual judgment a large share of the time, and a debrief on exactly what they had done wrong did not move the numbers on the next round. Vigilance is a property of how the decision is structured, not a property of operator willpower that can be moralized through training. What matters is how the AI's confidence is presented, what the operator is required to verify before acting on a recommendation, and what irreversibility gates sit downstream. Amazon's response to the AWS Kiro incident in February 2026 is the operational shape this recommendation prescribes: it added secondary-approval gates on irreversible actions, rather than a memo asking engineers to be more careful (Hale, 2026).

Invest in instrumented hold-both infrastructure.

The first three shifts imply an investment priority. The five clusters sit on the human side of the AI-human relationship, so their tensions surface there, and each cluster's design question targets one stage of cognitive passivity: calibration, attention, vigilance, skill, and shared understanding. The firm should spend less on tools that ask people to make a choice and more on tools that help two real goods coexist, instrumented so the firm can watch how each tension evolves rather than declaring it resolved when it is not. The dilemma-collapse condition adds one piece of discipline to this investment. Forced-choice tooling is the right instrument only where a tension's verification surface is narrow, and its failure is bounded, which here is the rare exception rather than the rule. Outside that narrow pocket, a forced-choice framework misdiagnoses the tension and stalls the decision by turning it into a debate over which side to pick.

V · Discussion · 5.3

Theoretical embedding

The findings sit within two established theoretical lineages and extend each by anchoring it in the agentic era.

The first lineage is the foundational human-automation canon, running from the levels-of-automation taxonomy (Sheridan et al., 1978) through Bainbridge's irony of automation (Bainbridge, 1983), Endsley's situation-awareness and out-of-the-loop work (Endsley, 1995; Kaber & Endsley, 1997), Parasuraman and Riley on automation misuse and disuse (Parasuraman & Riley, 1997; Parasuraman et al., 2000), and Lee and See on trust calibration (Lee & See, 2004).

The canon was built for deterministic automation, across decades when computer systems carried out narrowly specified tasks under predictable conditions. Its substantive findings carried straight into the agentic era. The systems most likely to breed complacency are the ones that work best. Asking a disengaged operator to recover from a rare failure is psychologically untenable. Operator trust drifts away from system reliability, producing overtrust that shows up as misuse and undertrust that shows up as disuse. This paper maps each cluster onto one of these canonical strands: Trust Miscalibration onto Lee and See; Automation Bias and Complacency onto Parasuraman and Endsley; Skill Atrophy onto Bainbridge; and Cognitive Debt onto the later extension of the same logic to team understanding.

The thesis extends the canon in three specific ways. It retests the canon's mechanisms against contemporary evidence and finds they still hold: in radiologists missing the cancers the AI missed (Lyell & Coiera, 2017), in autonomous coding agent incidents in 2026, and in EEG-measured engagement deficits (Kosmyna et al., 2025). It converts the canon's loose-tension language into a formal dilemma-versus-paradox classification, which allows findings to accumulate comparably across studies in a way the canon's narrative format never allowed. And it reads the same mechanisms as a single connected cascade: Trust Miscalibration leading to Automation Bias leading to Complacency leading to Skill Atrophy leading to Cognitive Debt, rather than as five separate observations.

The second lineage is Smith and Lewis's (2011) Dynamic Equilibrium Model of paradox, which supplies the dilemma-versus-paradox distinction that the Dichotomy Probe puts to an evidential test. The model names four categories of paradox-generating activity, Learning, Belonging, Organizing, and Performing, and argues that organizational paradoxes are best managed by holding both poles in productive tension rather than settling them by forced choice. The five clusters do not all sit in one category. Three are Learning paradoxes, one is a Performing paradox (Automation Bias), and one is an Organizing paradox (Automation-Induced Complacency). Learning is the most common of the three. That fits a study whose subject is a single human learning to work alongside a capable system. The thesis extends the model in two ways relevant here. It supplies a portable diagnostic, the Dichotomy Probe, that turns the dilemma-paradox distinction from a conceptual claim into a procedure any team can apply against any evidence base with a traceable verdict at the end. And it supplies the Equilibrium Mapping, a visual rendering that makes the structure of each paradox auditable rather than merely argued, with the load-bearing pole on the left and the counter-pole on the right.

V · Discussion · 5.4

Critical appraisal

The findings rest on the scope and the corpus weight. Each carries limits worth stating plainly.

The first limit is scope. This paper covers only one supercluster of a wider thesis, Outsourcing Human Cognition. It sits entirely at the Human Tier and on a single cognitive layer, the operator working alongside an AI tool. See Appendix C.5.

The humanistic focus of this paper is also a boundary. A result drawn from one Tier cannot test whether the same paradox-dominance holds where the AI system is the subject or where the organization deploying AI at scale is the subject. The wider thesis frames a Tier-aligned affinity hypothesis: a supercluster's mechanism spine converges on the side of the AI-human relationship that matches its Tier. That hypothesis cannot be tested on a single supercluster because a single Tier provides only one data point. It supplies one confirming case, a Human-Tier supercluster converging on a Human-side spine, and nothing more. The other superclusters appear in this paper only as explicit out-of-scope boundaries and in future work.

The second limit is corpus selection bias. The 62 evidence items were assembled through targeted ingestion of foundational human-automation work and contemporary AI-era publications, with a depth-evaluation discipline that routed toward primary sources. That selection is not neutral. It over-represents English-language academic research and well-documented industry incidents, and it under-represents settings where these tensions surface but documentation is sparse: non-Anglophone regulatory contexts, frontier labor markets, and deployments kept commercially confidential. The paradox-dominance finding is robust within the in-scope evidence base, but it should not be read as a universal claim about every case of outsourcing cognition. A corpus assembled with a different sectoral mix might shift the distribution.

The third limit concerns the evidentiary weight of the corpus. Of the 62 pieces of evidence, 26 are academic papers, which is a 42 percent academic share. That share is uneven across the five clusters: 77 percent for Trust Miscalibration, 50 percent for Complacency, 38 percent for Skill Atrophy, 27 percent for Cognitive Debt, and 15 percent for Automation Bias.

VI · Conclusion · VI.1

Summary and Answering the Research Questions

This chapter closes the paper. It does two things. First, it summarizes the study on Outsourcing Human Cognition and directly answers the two research questions. Second, it looks forward: it names what a wider study would test, where future research can go, and what the findings mean for business leaders who hand cognition to agentic systems. It introduces no new analysis. It reports the verdicts reached by Chapter IV and answers the questions posed by Chapter I. Where a claim rests on specific evidence, this chapter points back to Chapters III and IV rather than re-citing the lead sources.

Every AI breakthrough arrives with its yet. Agentic performance improves, yet employee trust lags behind or dangerously races ahead of what AI has actually earned. Models become more capable, yet individuals exercise their own judgment less and less. Reliability increases, yet that same reliability draws staff into overreliance. Automation spreads, yet teams expected to intervene when AI fails are less able to do so. Output scales, yet the organization understands less of what it ships.

The paper studied one supercluster, Outsourcing Human Cognition, and its five recurring tensions. The work read a corpus of evidence, identified the tensions that recur across it, and ran each tension through the seven-step Dichotomy Probe described in Chapter III.

The Probe reframes a tension as a yes-or-no question, tests both answers against the evidence, and sorts the tension into one of two kinds: a dilemma, also referred to as a trade-off (a forced either-or choice) or a paradox, a non-competing choice (two poles held together). Chapter III produced the descriptive catalog. Chapter IV produced the verdict on each tension. Both research questions can now be answered directly.

Research Question 1:

What recurring business tensions arise from outsourcing human cognition to agentic AI? Chapter III gives the full description of each. Here is the executive summary:

  • Companies must choose between the speed of acting on agentic AI's confident output and the discipline of calibrating where that confidence is earned, so trust miscalibration doesn't push employees to rely on systems where they are least reliable. Either way, the company pays: overtrust lets polished errors flow into real decisions; undertrust leaves a working system underutilized.
  • Companies must choose between optimizing agentic AI for maximum speed and scale, and slowing down to embed structured verification and guardrails against automation bias, so employees don't silently stop questioning what the system produces. Either way, the company pays: trust the outputs and rare errors quietly flow into real decisions; verify every output, and the speed gained evaporates.
  • Companies must choose between running a near-flawless agentic system unattended, and holding employees to active monitoring against automation-induced complacency, so the system’s dependability doesn't erode the habit of watching it. Either way, the company pays: look away, and the rare error surfaces only as a full-blown incident; demand constant vigilance over a system that rarely fails, and you ask for something no team can sustain.
  • Companies must choose between optimizing agentic AI for maximum efficiency on work humans once performed themselves, and rotating those humans back through the task to hold off automation-induced skill atrophy so their competence doesn’t drain away through disuse. Either way, the company pays: lean fully on the agentic system, and the day it fails, no one remembers how the work was done; preserve the skill, and you forfeit some of the savings the automation promised.
  • Companies must choose between optimizing agentic AI for fast, finished deliverables and making teams do enough of the work themselves to avoid accumulating cognitive debt, so throughput doesn't outrun the understanding that keeps a system maintainable. Either way, the company pays: ship at the system's pace, and the team's grasp of its own work quietly erodes; slow down to build that shared understanding, and the productivity gain shrinks.

Research Question 2:

Which of these tensions are dilemmas? (i.e., either-or, trade-offs) And which are paradoxes? (i.e., both-and, non-competing choices)

All five tensions are paradoxes. None resolves into a single defensible pole. Each holds two load-bearing poles, and none of them are defensible based on research evidence. This chapter states the verdict; it does not rerun the Probe.

The basis for each verdict, in one line and drawn from Chapter IV, is as follows. For Trust Miscalibration, both overtrust and undertrust are costly, and the tension persists through training and skill, so it must be held in balance rather than chosen between. For Automation Bias, speed and scrutiny compete at the moment of decision: scrutinizing every output erases the speed gain, and scrutinizing none lets the error through. For Automation-Induced Complacency, the benefit arrives now, and the cost arrives later, so the usual fixes of training and incentives do not dissolve the tension. For Automation-Induced Skill Atrophy, heavy automation and regular manual practice are both legitimate, and the trade-off must be managed over time rather than settled once. For Cognitive Debt Accumulation, the same workflows that make a team feel faster erode the shared understanding that keeps the work safe to change.

The agentic paradox

The paper's title promises business dilemmas in outsourcing human cognition. The finding inverts the everyday word. What practitioners reach for as dilemmas, forced either-or choices, are, on inspection, paradoxes in disguise. As discussed in the introduction, business literature has almost preconditioned leaders to think in terms of trade-offs: long-term versus short-term, exploration versus exploitation, speed versus security.

But in the context of outsourcing cognition to agentic AI, those familiar dichotomies no longer hold. What looks like a dilemma, i.e., push harder on one pole, accept a cost on the other, reveals itself as a paradox, where the attempted solution quietly deepens the very vulnerability it was meant to offset.

Leaders are not just deciding where to take the loss; they are deciding how much to rely on the system, and that reliance erodes the very judgment their plans depend on. The choice is not a single cost paid once and done. It is a choice about which human ability they let fade, and each option eats away at the ability it depends on. Trust the system or check it; automate or keep practicing; move fast or stay careful. Each looks like a clean trade: pick a side, accept a cost on the other. But the cost is neither fixed nor separate.

The side a leader leans on slowly erodes the human ability they would need to switch back or step in when the system fails. Trust without checking wears down the judgment that would have caught the system’s mistakes. Automation without practice drains the skill that stepping in requires. Speed without care thins the understanding that keeps the work safe to change. Each time, the fix deepens the very weakness it was meant to offset. The evidence shows that every apparent choice is really a pair of poles that must be held together because neither side protects what the other needs. Taking only one side leaves the leader with no defensible position.

This is the agentic paradox. Handing cognition to an agentic system does not remove human judgment from the loop. It relocates the tension. The work shifts from doing the task to maintaining a balance between relying on the system and being able to control it.

The five tensions share one structure, and naming it is the synthesis this chapter delivers. Each cluster is a paradox to be managed, not solved. Each one draws on a single scarce human faculty pulled in two directions at once.

Trust Miscalibration draws on an accurate read of how reliable the system is for the task. Automation Bias draws on critical scrutiny at the moment of decision. Automation-Induced Complacency requires sustained vigilance. Automation-Induced Skill Atrophy draws on practice, the time a person spends doing the work by hand. Cognitive Debt Accumulation draws on shared understanding. In every case, the faculty cannot be fully given to the machine and fully kept by the human at the same time. That shared structure is why these are paradoxes, not dilemmas. Outsourcing cognition does not remove the tension. It relocates it.

One question remains: why this pattern, and why does it recur? There may be a structural reason the five tensions all take the form of paradoxes, introduced in Chapter II and worth recalling here. Frontier AI has a jagged capability surface (Karpathy, 2024): it is superhuman on some tasks and surprisingly weak on nearby trivial ones, and humans cannot tell in advance which is which. A dilemma needs one pole that is safe to choose. A jagged substrate denies that safety on both sides at once. Trusting the system fails because the gaps bite without warning. Refusing to trust it fails too, because walking away from the tasks it does well costs more than the errors it avoids. Neither pole survives alone, so the tension takes a paradoxical shape by structural necessity.

Collective implications

Each agentic paradox in this paper has been analyzed individually, but they are best read together. And they point to something bigger. These five costs don’t show up one by one. As companies hand more judgment to agentic systems, calibration, scrutiny, vigilance, skill, and shared understanding all come under pressure at once and begin to compound. A team that has stopped checking the output of an agentic system is also a team that is losing the skill to check well. A team losing that skill is also a team whose shared understanding of their work may be thinning. What follows is atrophy: verification, vigilance, and skills all weakening together, enabled by the very reliability that makes agentic systems seem safe to deploy.

Under this pressure, human roles will evolve in more than one way. One change is obvious: the person is no longer the sole or primary owner of the output. More importantly, the role shifts from doing the work to holding the balance. The person becomes the one who decides how far to rely on the system and how much human capability to preserve. That is a different job, and it is harder to see because its product is a balance held rather than a task completed.

This raises a broader question. As outsourcing cognition scales across industries, who or what assumes responsibility for preserving collective human cognitive capability? If every firm allows the same faculties to atrophy at the same time, the capacity to recover when systems fail thins across entire fields, not just within any single company. That may be a question for another paper, but it is worth noting here.

VI · Conclusion · VI.2

Outlook

This paper is bound to one supercluster, Outsourcing Human Cognition. The other tensions of the agentic era, arising from governing algorithmic systems, deploying them at scale, and the societal implications of AI, were only noted as out of scope. The same methodology used to produce this paper can be extended beyond its limits.

Whether the paradox pattern holds elsewhere is exactly what that extension would test. The jagged-intelligence account predicts where it should break: where the capability gaps are bounded and the output is checkable, a tension should sharpen back into a real dilemma with one defensible pole. The governing and deploying domains likely contain more of those bounded-verification cases, so they may surface genuine dilemmas that this supercluster did not. A single clean counter-example would not weaken the framework. It would sharpen it by marking the line where paradox ends and forced choice begins.

Three directions for future research follow from the evidence. The first is longitudinal. Complacency and skill atrophy both rest on a delayed-cost mechanism, the benefit now and the bill later. Field studies that follow teams over time would test whether the managed balance actually holds or quietly tips once the costs come due. The second is measurement. Several of these tensions rest on an accurate read of how reliable a system is for a given task, and the paper treats that read as available. How an organization produces such a read in practice, and how it keeps the read current as the system changes, is open. The third is boundary-hunting. Future work should look for any tension in this supercluster that, under new evidence, resolves into a dilemma. The lone Automation Bias dilemma shows the Probe can return that verdict, so the question is empirical, not settled.

For practitioners, the finding changes the question, and that change is the point. A business leader who has taken the agentic paradox seriously stops asking which pole to pick. The question is no longer whether to trust the agentic system or to check it, whether to automate or to keep the human practice alive. A paradox cannot be solved by choosing a side, because neither side is safe to choose. So the useful question becomes a design question: how might we hold both poles, so that the system’s strengths and the human’s strengths check each other rather than erode each other?

Chapter IV phrases that question for each of the five tensions. This chapter only reminds the reader that those questions exist and that they are where the work goes next.

That shift is the paper's hand-off to design. A How-Might-We question is not a loose end left after the analysis. It is a deliberate move that opens a space to act in. It treats a hard, irreducible tension with an open mind rather than a forced choice, and it invites the business leader, the designer, and the researcher to build inside the space rather than to settle the question once.

This paper diagnoses. It shows how the business tensions that emerge when organizations outsource human cognition to agentic systems are paradoxes to be managed, not trade-offs to be resolved. It maps each one, and the scarce capacity it spends in two directions at once. The diagnosis is where the analysis ends, and the design begins.