MBA Thesis  •  Chapter III

Methodology and Results

This chapter explains how the paper addresses its two research questions and reports the data collected. It has four parts. Section 3.1 justifies the method by showing why the two questions need each other, and why three obvious alternatives each fail one of them. Section 3.2 sets out the research design: the four-database evidence spine, the seven-step Dichotomy Probe that classifies each tension, and the Equilibrium Mapping that renders the paradox verdicts. Section 3.3 describes how the corpus was built, source by source and tension by tension, and how each source was graded for quality. Section 3.4 reports the result of that collection: a plain description of the five recurring tensions in outsourcing human cognition, each with its definition, its business tension, its evidence base, and its two pole pairs. The verdicts themselves, dilemma or paradox, belong to the analysis in Chapter IV, not here.

5
recurring tensions in the Outsourcing Human Cognition supercluster
62
evidence entries across the five clusters, each tied to one source
42%
of evidence entries drawn from academic papers
61/62
evidence instances classifying as paradoxes among documented capability-cost tensions

3.1 · Justification of the selected method

Two questions that need each other

The two research questions pull in different directions, and neither can be answered without the other. RQ1 asks which recurring business tensions arise from outsourcing human cognition to agentic AI. That is an empirical question, and an adequate answer is a catalog grounded in primary evidence. RQ2 asks which of those tensions are trade-offs, a choice between either one pole or the other, and which are non-competing choices that hold both poles at once. That is a classification question, and an adequate answer is a procedure that can be applied to each tension on its own and returns a verdict that another researcher could reproduce. A catalog without a classification tells a manager nothing about how to handle any particular tension. A classification scheme with no catalog has nothing to classify. The method therefore has to deliver both an evidence base and a repeatable way of sorting that evidence.

Three obvious alternatives were considered and set aside, because each answers one research question well and the other not at all.

SET ASIDE 01

A pure literature review

A pure literature review could catalog the recurring tensions. The foundational human-automation canon already describes their structure in depth: Sheridan et al. (1978), Bainbridge (1983), Endsley (1995), Parasuraman and Riley (1997), Parasuraman, Sheridan, and Wickens (2000), and Lee and See (2004). What that canon cannot do is return a verdict on a specific tension a decision-maker faces today. It describes the tensions, but it does not say whether the right response in a given 2026 deployment is to choose between the poles or to hold them together. RQ2 is exactly that operational question.

SET ASIDE 02

A practitioner survey

A practitioner survey could produce classification verdicts at scale, but the verdicts would be opinions about the tensions rather than tests of them. Someone who has built a business around treating a tension as a forced choice will classify it as one. Someone who has built a business around holding it open will classify it as a paradox. The instrument would record where the field's intuitions sit, not where the structure of the tensions sits, and RQ1's requirement of grounding in primary evidence would fail.

SET ASIDE 03

A single deep case study

A single deep case study could ground a classification in primary evidence, but it could not establish a pattern. The empirical claim of this paper is that outsourcing human cognition takes a recurring form, and that the same kind of tension surfaces across different settings and stakeholder lenses. One case shows the phenomenon vividly. It cannot show that the phenomenon recurs.

The method adopted here is the Dichotomy Probe paired with Equilibrium Mapping, applying tension by tension across an assembled corpus. It answers both questions from one evidence base. The corpus answers RQ1: it surfaces, names, and locates the recurring tensions. The Probe answers RQ2: it tests each tension against documented evidence, and it returns a trade-off-or-paradox verdict. Because both rest on the same primary sources, the catalog and the classifications cannot drift apart. Each verdict leaves a written trail another reader can follow. At the level of the whole corpus, the distribution of verdicts is what RQ2 is asking for.

One caveat belongs here rather than in the limitations chapter because it shapes how the headline result should be read. The corpus was assembled by sampling documented capability-cost tensions, the pattern set out in Chapter I in which every breakthrough arrives with its own cost. That sampling frame already leans toward cases where both poles are real, which is close to the structural signature of a paradox. The finding reported in Section 3.4 and analyzed in Chapter IV is therefore stated conditionally: among documented capability-cost tensions in outsourcing human cognition, 61 of 62 evidence instances classify as paradoxes. It is not a claim about all business tensions everywhere. The single dilemma in the corpus matters for exactly this reason. It sits within the Automation Bias cluster, indicating that the Probe can still return the other verdict, so the instrument is not simply restating its own input.

3.2 · Research design

Four working parts, two core instruments

The design has four working parts. The four-database evidence spine assembles and structures the corpus. The Dichotomy Probe produces each verdict. The Equilibrium Mapping renders the paradox verdicts. A Tier-aligned affinity rule governs how clusters gather into superclusters, and a naming exception lets long-established research terms stand. The Probe and the Mapping are the core instruments. The spine feeds them, and the affinity rule and the naming exception govern how individual verdicts aggregate into the catalog's structure.

The four-database evidence spine

The corpus is held in four linked databases that run from raw material to grouped findings. Sources come first. Each source is a self-contained archive of one cited item. Each source is then converted into one or more Evidence entries. An evidence entry is a single tension instance extracted and coded from a source. Evidence entries gather into Clusters. A cluster is one recurring tension, defined by its own evidence base and carrying the verdict. Clusters gather into Superclusters, grouped by where their tensions operate. This paper covers one supercluster, Outsourcing Human Cognition, and its five clusters. The spine matters for the method because it keeps every claim traceable to a primary source and ensures that the catalog (RQ1) and the classification (RQ2) rest on the same records.

The recurring tensions were not imposed from the top down. They emerged from the corpus. As each source was evaluated, the tension it documented was abstracted from the bottom up into a controlled vocabulary of canonical, recurring tensions. Each canonical tension was defined by synthesizing its own evidence base, tagged with a paradox category after Smith and Lewis (2011), and assigned to a cluster and a supercluster. Each canonical tension was then curated through a lifecycle of three states: proposed, stable, and deprecated. This inductive distillation is itself part of the method. The five tensions reported in Section 3.4 were stress-tested against the corpus rather than presupposed.

The Dichotomy Probe

The Probe tests whether a candidate tension is a trade-off, a forced choice between mutually exclusive options, or a paradox, a persistent tension between interdependent goods. It is binary in form, a yes-or-no at each step, but evidential in its verdict, graded against documented sources rather than intuition. The seven steps run in order.

#StepWhat it does
1FrameIdentify the tension as it surfaces in the evidence, and name the business domain in which it manifests.
2Build the questionWrite a single yes-or-no sentence of the form "In this business domain, does this assertion about one pole hold?" The question must be anchored to a domain, neutral between the poles, and open to being proven false.
3Run the rebuttalsBuild the defense of the yes pole, citing documented cases where the antecedent produced the consequent. Then build the defense of the no-pole, citing documented cases where it did not. Both defenses must draw on the corpus rather than assert against it.
4Decompose into sidesName an AI-side pole pair and a human-side pole pair, each written as two short nouns held in tension.
5Independence testHold each side's design dial at its most favorable setting, and check whether the other side's tension still has a non-zero base rate. If it does, the two sides are independent counterforces. If it does not, the framing has one axis expressed twice, and the question is re-formed.
6Assign load-bearing polesOn each side, the pole whose neglect produces the failure the deployment is trying to prevent is the load-bearing pole. The assignment is defended in one written sentence per side.
7Author the mapping and the design questionThis step runs for paradoxes only.

Verdicts are rendered in three ways.

VERDICT · TRADE-OFF

A forced choice

A trade-off appears when Step 3 yields asymmetric defenses, one side defensible and the other not. The Probe stops there and routes to choice logic.

VERDICT · PARADOX

Interdependent goods

A paradox appears when both sides are indefensible alone, when Step 5 confirms independence, and when Step 6 finds a defensible load-bearing pole on each side.

VERDICT · FALSE DILEMMA

A broken question

A false dilemma is a transient result that fires when the framing is broken, for example, when both sides hold up, or when independence fails twice. A false dilemma is never stored as a verdict. It triggers a reformulation and a re-run. That third outcome is what allows the Probe to expose a broken question rather than force it into a category.

A worked case shows Step 5 in motion, using the Trust Miscalibration cluster. The AI-side dial governs how clearly the system shows its own reliability. The human-side dial governs how the operator sets reliance: informed reliance at one end, miscalibrated trust at the other. Hold the AI side at its most favorable setting, a system that shows its confidence and the basis for it as clearly as any deployed system does, and ask whether miscalibrated trust still has a non-zero base rate. It does. Becker and colleagues (2025) measured experienced developers’ misjudging of an AI assistant's effect on their own speed by roughly forty points, even with the tool in hand. Now hold the human side at its most favorable setting, an operator calibrating reliance carefully against observed behavior, and ask whether opacity still bites. It does. A careful operator cannot calibrate against a failure surface that the system does not expose. Because neither side's best case drives the other's base rate to zero, the two are independent counterforces rather than one axis written twice, and the tension routes toward a paradox rather than a re-formed question.

This paper reports the procedure, not its verdicts. Chapter IV runs the Probe to a verdict for each of the five clusters. The paradox category it assigns each cluster comes from the Dynamic Equilibrium Model introduced in Chapter II, which this chapter applies rather than re-teaches.

Equilibrium Mapping

For paradoxes only, the verdict is rendered as a five-column mapping. The AI system sits on one side and the human factor on the other. Each arm carries a load-bearing pole and its counter-pole, with the load-bearing pole always on the left. The organizing image is Alexander Calder's mobile: a horizontal rod holds two arms in counterbalance; each arm maintains its own internal tension; and the paradox is the balance of the whole. Disturb one pole and the mobile re-settles. No single pole wins. Trade-offs are not mapped this way because the question they pose, which pole is defensible, differs in kind from the paradox question of how the two arms balance. The load-bearing assignment is made for each instance of tension, not once per cluster. Two evidence entries that share a pole pattern at the cluster level can still have their poles weighted differently, because the weighting depends on the stakeholder lens and the business domain documented in each case. In this paper, the Equilibrium Mapping is a method artifact, described here and applied in Chapter IV. Section 3.4 reports each cluster's two pole pairs as a plain description, without mapping them to a verdict.

The Tier-aligned affinity rule

Clusters gather into superclusters, and each supercluster carries a Tier that names where its tensions operate: Human, for failures within the operator; System, for mismatches between the AI and the workflow; Organization, for governance and process pathologies; and Society, for tensions held outside the firm's actionable scope. The supercluster studied here, Outsourcing Human Cognition, carries the Human Tier. All five of its clusters sit inside the operator's own cognition. The affinity rule proposes that a supercluster's evidence converges toward the AI-human pairing that matches its Tier. A Human-Tier supercluster converges on the human side, while the AI side fragments across several surfaces. The rule is an organizing hypothesis rather than an established result, and a single Human-Tier supercluster cannot test it on its own.

The literature-anchored naming exception

The default rule names each cluster after its AI-side load-bearing pole. The exception keeps the established research term where one already covers both poles of a tension. The five clusters here, Trust Miscalibration, Automation Bias, Automation-Induced Complacency, Automation-Induced Skill Atrophy, and Cognitive Debt Accumulation, keep their canonical names rather than being relabeled. Decades of human-automation research by Sheridan, Bainbridge, Endsley, Parasuraman, Lee, and See have developed a recognized vocabulary that encompasses all of these tensions. Renaming Automation Bias to its load-bearing pole would shed that anchor without adding clarity. The exception is invoked explicitly where it applies, and new clusters default to the pole-concept rule.

3.3 · Data collection

Sources first, tensions second

The corpus was built in two stages, sources first and tensions second, each with its own protocol and quality checks, and each producing records that survive after the original web address goes dead or the publishing platform changes.

Sources

Every source enters through a source database and is archived as a self-contained record: a full summary written from the primary material, the converted full text or a transcript, a resolvable APA reference, and a local copy of any document that does not convert cleanly to text. The governing test is whether a committee member, working offline, could verify a cited claim from the record alone. The second discipline of source ingestion is a depth-evaluation pass. When a lead was itself secondary, a news article summarizing a study, a podcast referencing a paper the guest had recently published, or a business report citing an underlying dataset, the protocol required retrieving the primary source and treating it as the citable ground truth, keeping the secondary item only as accessible framing. The effect is that the citation chain ends at primary research wherever the original lead pointed to coverage of it. That is what makes the recent AI-era findings defensible against the charge that they rest on journalism rather than on research.

Leads were gathered as a funnel rather than a single search, from three source pools. The first pool consists of peer-reviewed academic work, identified backward from the foundational human-automation canon and its citation network. The second pool is practitioner and industry material, found forward from contemporary AI-era coverage in the trade and research press between 2024 and 2026. The third pool is public case documents and regulatory material, found laterally from the reference lists of sources already admitted. A lead was admitted as a source when it met three tests: it documented a concrete capability-cost tension rather than asserting one, it resolved to a primary record that an offline reader could verify, and it fit the supercluster's scope. A lead was excluded when it only restated a tension already covered by a stronger primary source, or when its claim could not be traced past secondary coverage. Screening for a cluster stopped when new leads no longer surfaced new tension shapes, indicating the point at which the cluster had reached saturation.

Each source carries a Source Type drawn from a typology of fourteen types, tiered by evidential weight. The academic tier holds Academic Paper, Academic Report, Research Study, and Book. The practitioner tier holds Article, Interview, Survey, Business Report, and Conference Talk. The primary-media tier holds Video, Documentary, and Excel data. The lowest tier holds blog posts and blog comments. This typology is the operational basis for the quality assessment and for the academic-share figures reported below.

Source quality was assessed against a six-part framework, the 6Rs: realness, richness, repetition, rationale, repartee, and regal standing. Format is not the same as independence, and the corpus has to be read with that distinction in view. Several of the most quotable figures in this field come from the weaker end of the quality gradient and from parties with an interest in the result, such as vendor studies of their own products or executive statements about internal adoption. Where such a statistic carries an argument, it is corroborated with an independent source or flagged in the text. Chapter V returns to this gradient as a limit on the weight any single figure can bear.

Tensions

Each recurring tension extracted from a source is recorded as one evidence entry. It is classified into one category, trade-off or paradox, and it is tied to one business domain, one cluster, a level-of-automation interval after Sheridan, and a paradox category after Smith and Lewis. The load-bearing feature of this stage is that classification precedes template selection. The Dichotomy Probe runs first, as a blocking gate, and only once the verdict lands is the matching template filled in. Trade-offs take a decision template, with poles, decision criteria, trade-off magnitude, reversibility, and regret. Paradoxes take a Calder-mobile template, the Equilibrium Mapping, and its paired How We Might question. The order matters because an earlier version of the tool filled in a paradox template regardless of the verdict, so it recorded paradoxes and trade-offs alike as paradoxes by default. Running the classification first makes that particular error impossible. A consistency check runs before every entry is committed. A trade-off entry may carry no Equilibrium Mapping. A paradox entry must carry one. Every entry must link to exactly one source, one cluster, and one canonical business domain. Any failure aborts the write rather than saving a partial record.

This source-to-evidence conversion is what turns a catalog of readings into a gradeable dataset. It is also what lets the academic-share figures below be read directly from the corpus rather than estimated.

Corpus description

The supercluster studied here holds five clusters and sixty-two evidence entries, distributed thirteen, thirteen, twelve, thirteen, and eleven across the five clusters in order. Twenty-six of the sixty-two evidence entries are drawn from academic papers, representing a forty-two percent share. That share varies sharply by cluster, as the descriptions below make clear: Trust Miscalibration draws 77% of its evidence from academic papers, Automation Bias 15%, Automation-Induced Complacency 50%, Automation-Induced Skill Atrophy 38%, and Cognitive Debt Accumulation 27%. The variation tracks how recent each tension is. The older, lab-grounded tensions rest on academic work, while the newest tensions rest more on practitioner reports and field accounts, because the academic literature has not yet caught up to them.

3.4 · Results of data collection

The five recurring tensions the corpus surfaced

This section describes the five recurring tensions that the corpus surfaced, in order: Trust Miscalibration, Automation Bias, Automation-Induced Complacency, Automation-Induced Skill Atrophy, and Cognitive Debt Accumulation. Each tension is given a definition, a statement of the business tension it creates, an account of its evidence base, and its two pole pairs, one on the AI side and one on the human side. The account here is descriptive. It reports what each tension is and what evidence supports it. It does not classify the tension as a trade-off or a paradox, and it does not fill the Equilibrium Mapping with a verdict. Those steps belong to the analysis in Chapter IV.

3.4 · Results · Cluster 1 of 5

Trust Miscalibration

Trust miscalibration is the gap between the confidence humans place in agentic systems and what those systems actually deliver. That gap is worth guarding against because it can lead people to overtrust the systems where they are least reliable and to distrust them where they perform best. Trust can also be driven by how an agentic system presents itself, by its fluency and its veneer, rather than by what it can actually do.

Business Impact

An agentic system's smooth, well-phrased output invites trust, whether or not the answer is correct. That polish is learned during training. It is not evidence of the system's ability to solve the task at hand. When employees act on that impression and lean on an agentic system beyond its competence, silent failures slip through. At the other extreme sits undertrust, which a single visible error is often enough to trigger. When a critical mistake surfaces, or after a few disappointing ones, users may pull back and leave the system underused. Both reactions are expensive for firms. Overtrust lets bad outputs into real decisions, and even when those outputs are caught, the review and rework needed to catch them can quietly erase the speed gains the agentic system promised. Undertrust, meanwhile, may leave a working agentic system sitting on the shelf.

Supporting Evidence

The research behind this cluster is unusually deep for a problem this recent. It draws on ten sources and thirteen evidence entries that range from controlled lab experiments and a randomized field trial to two foundational reviews of the human-automation literature. Its anchor is Lee and See (2004), whose review of trust in automation remains the authoritative reference for what calibrated reliance entails: trusting the system in line with what it can actually do. The cluster sits inside the canon that Sheridan, Bainbridge, and Parasuraman built, in which the danger was never the machine alone, but the way human judgment bends around it.

Lee and See (2004) set out the core mechanism in plain terms. People misuse a system when they trust it more than its record warrants, and they stop using a system when they trust it less than it deserves. They named two failure cases on opposite sides. The cruise ship Royal Majesty drifted off course for a full day and ran aground while the crew leaned on a faulty navigation system. An Airbus A320 crew failed to take manual control as the autopilot flew them into the ground. Hoff and Bashir (2015) build directly on that work in a review of 127 studies. They show that people tend to assume a machine is near-perfect at the start, so early trust runs ahead of real reliability. Two more fatal cases mark both errors: the Costa Concordia, after a captain ignored the navigation system, and a Turkish Airlines crash, after pilots kept trusting an autopilot fed by a broken sensor.

The other sources show the same pattern across different settings, starting with individual users. Bainbridge and colleagues (2011) found that a robot's physical presence alone, not its skill, pushed people to obey an odd, destructive order. Glikson and Woolley (2020) report in their review that most participants kept following a human-like robot even as it made plainly unreasonable requests. Klingbeil and colleagues (2024) showed that simply labeling advice as coming from an AI made people follow it against clear evidence and against their own stated belief. Shneiderman (2020) traces how human-like design invites users to expect more competence, and more responsibility, than the machine actually has.

The pattern repeats with engineering leaders and developers. The Cortex (2026) benchmark found engineering leaders convinced that AI had helped, while their own metrics moved the wrong way. Two readings of the same field trial sharpen the point. Osmani (2025) and Becker and colleagues (2025) describe how experienced developers stayed confident that AI sped them up by about twenty percent, after timed records showed it slowed them by nineteen percent. The slowdown came less from bad code than from the time spent prompting, waiting, and reviewing, a quiet tax that swallowed the speed the tool seemed to offer. Read from the other side, the sharpest case of undertrust comes from a single study. Dietvorst and colleagues (2015) found that people dropped an accurate forecasting model after one visible mistake, and chose a worse human forecaster they had already watched do worse.

SourceDescriptionInsight
Lee and See (2004)A seminal review that crystallized the terms appropriate reliance and trust calibration. It defines two failure modes: misuse, when people rely on a system more than its record warrants, and disuse, when they reject a capable system. It documents the Royal Majesty running aground after a day adrift, and an Airbus A320 crew that failed to take manual control.People rely on a system in line with how much they trust it, so trust set wrong leads to misuse or disuse.
Hoff and Bashir (2015)A review of 127 studies finds that people misuse automation when they over-trust it and stop using it when they under-trust it. Two cautionary cases: the 2012 Costa Concordia sinking after the captain ignored the navigation system, and the 2009 Turkish Airlines Flight 1951 crash after pilots kept relying on an autopilot fed bad data by a failed sensor.People often assume a machine is perfect at first, so they trust it more than its real reliability warrants.
Klingbeil and colleagues (2024)In an incentivized trust-game experiment with 319 UK participants, labeling identical advice as AI rather than expert raised cooperation to 65 percent versus 9 percent in the control group when the partner had mostly defected, and following AI advice cut the advisee's payoffs 22 percent below control.People follow advice more once it is labeled AI, even against clear evidence and their own judgment.
Shneiderman (2020)A commentary naming concrete harms from excessive automation: the Boeing 737 MAX crashes, stock-market flash crashes, and parole or hiring decisions made from machine learning trained on biased data. It warns that designing robots to look and talk like people invites three trust errors.Human-like robot design makes users expect more responsibility and competence than the machine actually has.
Glikson and Woolley (2020)A 20-year review of about 150 empirical studies finds miscalibration runs both ways: low trust in a capable system causes disuse, high trust in a weak system causes misuse. In one cited study, 91 percent kept following a human-like robot after watching it err.When a system looks more capable than it is, people follow its bad outputs and stop checking.
Bainbridge and colleagues (2011)In a book-moving study with a humanoid robot (59 usable participants), 12 of 20 people obeyed an odd instruction to throw textbooks in the garbage when the robot was in the room, versus 2 of 20 over live video and 3 of 19 over augmented video.A robot's physical presence alone, not its competence, made people obey an odd, harmful instruction.
Cortex (2026)In a 2026 engineering benchmark (about 50 engineering leaders), 91 percent said AI improved developer velocity and quality, but only 25 percent had data to back that up, while measured indicators moved the wrong way: incidents per pull request rose 23.5 percent and change-failure rates climbed roughly 30 percent.Engineering leaders believe AI helped far more than their own metrics actually show.
Osmani (2025)Reports on a randomized controlled trial of 16 experienced open-source developers working on their own large codebases; using AI tools made them 19 percent slower on average. The same developers expected a speed-up of about 24 percent beforehand, and afterward still believed AI had made them about 20 percent faster.Developers kept believing AI sped them up even after timed records showed it slowed them down.
Becker and colleagues (2025)In a randomized trial, 16 experienced open-source developers worked 246 real issues on repositories they had averaged five years on; allowing AI tools made them 19 percent slower, yet they estimated a 20 percent speed-up afterward and forecast 24 percent beforehand. They accepted under 44 percent of AI suggestions and made major changes to 56 percent of accepted ones.The slowdown came less from bad code than from the prompting, waiting, and reviewing tax that swallowed the tool's speed.
Dietvorst and colleagues (2015)Five incentivized studies establish algorithm aversion. People forecast real outcomes and could pick an algorithm or a human forecaster. After watching the algorithm make a mistake, they dropped it for a human forecaster they had already seen do worse on average.People drop an accurate model the moment it slips once, and keep a human forecaster they know is worse.

Business Tension

Companies must choose between the speed of acting on agentic AI's confident output or the discipline of calibrating where that confidence is earned, so trust miscalibration doesn't push employees to rely on systems where they are least reliable. Either way, the company pays: overtrust lets polished errors flow into real decisions; undertrust leaves a working system underutilized.

AI-side tensionHuman-side tension
On-task reliability gap versus Capability-grounded transparencyMiscalibrated trust versus Informed reliance

3.4 · Results · Cluster 2 of 5

Automation Bias

Automation bias is the human tendency to defer to automated or agentic systems over one's own judgment, accepting their outputs with less scrutiny than one would apply to one's own work. It is a deference at a point in time: at the moment of decision, the more capable the agentic system appears, the more readily its output is taken as correct.

Business Impact

Checking agentic outputs is cheap when structured tests exist, and costly when quality is a matter of judgment. Either way, the same speed and scale that make an agentic system worth adopting also make its output hard to verify, because it generates more good work, faster, than any human can reliably verify by hand. Reliability has a second consequence too: the better the agentic system performs, the less anyone feels the need to double-check its output. So companies face a difficult decision. If their employees verify each output, much of the speed and scale the agentic system promised can be eaten up by the checking itself. If they stop checking, the system's rare errors flow straight into real decisions. Those errors are rare, but they are more costly to the company, because they tend to surface only after they have already caused real damage.

Supporting Evidence

The founding evidence for this tension comes from medicine. Lyell and Coiera (2017) reviewed forty studies of clinicians using automated decision support, and they found a counterintuitive result: the more accurate the automation, the more often people deferred to it without checking. That finding belongs to a longer human-automation tradition. It runs from Bainbridge's work on the ironies of automation, through Parasuraman's studies of complacency, to Endsley's account of operators drifting out of the loop. The supporting research here spans a controlled monetary experiment, a reasoning benchmark, a field survey of nearly 29,000 developers, software-delivery telemetry, and documented production incidents. Each setting tests one claim: capable agentic systems invite the kind of trust that quietly suspends independent checking.

The strongest evidence for this tension is also the most controlled. Klingbeil and colleagues (2024) ran an incentivized experiment with 319 participants, real payoffs, and a third party who could be harmed by a bad call. They varied one thing only, a label, presenting identical advice as coming from an AI rather than a human expert. That single word moved behavior. About 44 percent of participants overrode their own stated assessment, payoffs dropped by roughly a fifth, and the expert-labeled version produced no such effect. Training participants beforehand did not fix it. The study isolates the core of the problem: trust follows the AI label, not any proof that the system is right.

The next question is why such misplaced trust survives contact with a wrong answer. Shen (2025) answers it with a benchmark of 2,686 mathematical problems, each stripped of one necessary condition. Models supplied an answer anyway in more than 90 percent of cases, and they were zero percent accurate when they silently assumed the missing piece. Asked to seek the condition first, they recovered it 91 to 97 percent of the time. The lesson is that a confident, fluent output gives the operator no signal that verification is required. That is what makes the failure silent.

Lyell and Coiera (2017) provide the founding field study and the mechanism that ties the set together. Reviewing forty studies of clinical decision support, they found automation bias in 81 percent of omission studies and 91 percent of commission studies, and, against intuition, more bias as accuracy rose. Their decisive contribution is naming verification complexity, the effort it takes to confirm a recommendation, as a better predictor of error than how many tasks the operator juggled. Their clinical data ground the cost in patients: computerized heart-trace misdiagnoses went uncorrected about 24 percent of the time.

The pattern then moves into knowledge work and operations. Stack Overflow's 2025 survey of 28,930 developers shows the benefit at field scale. About 70 percent report faster task completion, yet 87 percent remain concerned about accuracy, and 84 percent keep final human authority. That is vigilance and skepticism, rather than blanket complacency. Osmani (2025) supplies the cost ledger for software teams, drawing on a controlled trial and delivery telemetry across more than a thousand teams. Experienced developers ran 19 percent slower with AI while believing they were faster, and output rose alongside a 91 percent increase in review time and a larger bug load, with no gain at the organizational level. The hazard becomes concrete in Rosner and Uddin's (2026) account of an Amazon agent that deleted and rebuilt a production environment unprompted, causing a thirteen-hour outage that the remedy met with mandatory peer review. Ming (2026) widens the lens to general reasoning: in a forecasting study of 72 participants, 90 to 95 percent used AI as a substitute or as a source of confirmation, and only a small minority treated it as something to argue with.

SourceDescriptionInsight
Klingbeil and colleagues (2024)A controlled, incentivized experiment with 319 participants and real money. Labeling identical advice as coming from an AI rather than an expert drove about 44 percent of participants to act against their own stated judgment and cut payoffs by roughly a fifth; the expert-labeled text had no such effect, and prior training did not prevent it.The cleanest causal proof in the set: the AI label alone, with no demonstrated competence, induces over-reliance and imposes a measured cost.
Shen (2025)A benchmark of 2,686 reasoning problems with one necessary condition removed. Models answered anyway in over 90 percent of cases and were zero percent accurate when they silently assumed the missing piece, yet recovered the condition 91 to 97 percent of the time when prompted to ask.Explains why the failure is silent: a confident, fluent output is indistinguishable from a correct one, so the operator gets no signal that checking is needed.
Lyell and Coiera (2017)A systematic review of forty clinical decision-support studies. Automation bias appeared in 81 percent of omission and 91 percent of commission studies, higher-accuracy automation produced more bias, and verification complexity best predicted the lapse; computerized heart-trace misdiagnoses went uncorrected about 24 percent of the time.The founding source: it establishes the reliability paradox, names the verification-complexity mechanism, and supplies real clinical harm where deference reached the patient.
Osmani (2025)A synthesis of a controlled trial and delivery telemetry across more than a thousand teams. Experienced developers were 19 percent slower with AI while believing they were faster; output rose alongside a 91 percent increase in review time, larger changesets, more bugs per developer, and no organization-level gain.The single source carrying the cost-economics half: the same capability that lifts individual output shifts the bottleneck onto human review and pushes defects downstream.
Stack Overflow (2025)A field survey of 28,930 developers. About 70 percent report faster task completion and 69 percent report higher productivity, yet 87 percent remain concerned about accuracy, only 37.5 percent report better code quality, and 84 percent retain final human decision authority.Field-scale evidence of the genuine throughput benefit alongside pervasive skepticism and retained oversight.
Rosner and Uddin (2026)An investigation of a single incident at Amazon: an agentic coding tool autonomously deleted and rebuilt a production environment, causing a thirteen-hour outage. The agents held operator-equal permissions with no second approval; the remedy was mandatory peer review.A concrete, named-employer receipt that silent over-trust reaches customer-facing production, and that the corrective swing runs back toward human verification.
Ming (2026)A first-person account of a forecasting study with 72 participants. Most hybrid teams performed no better than AI alone: 90 to 95 percent used AI as a substitute or for confirmation, and only a 5 to 10 percent minority treated it as an adversarial sparring partner.Population-level evidence that deference is the dominant default, with a clean split between complacency and vigilance.

Business Tension

Companies must choose between optimizing agentic AI for maximum speed and scale, or slowing down to embed structured verification and guardrails against automation bias, so employees don't silently stop questioning what the system produces. Either way, the company pays: trust the outputs and rare errors quietly flow into real decisions; verify every output, and the speed gained evaporates.

AI-side tensionHuman-side tension
Verification overload versus Decision accuracyCognitive passivity versus Operator vigilance

3.4 · Results · Cluster 3 of 5

Automation-Induced Complacency

An agentic system that almost never fails is the hardest to keep watching: the more dependable it proves, the less reason humans have to check it. Automation-induced complacency is the tendency to lower active monitoring of an automated system that has proved dependable, sliding from close supervision to a passive sign-off given with barely a glance. Where automation bias is about wrong or missed actions after employees followed poor advice, automation-induced complacency is about slow, weak, or absent responses, because employees were not watching.

Business Impact

An agentic system that fails often keeps employees sharp, but one that almost never fails gives attention nothing to catch, so the habit of watching quietly fades. Demanding constant vigilance over an agentic system that rarely errs is an effort no team sustains, and the reliability that makes the automation worth running tends to breed the inattention that undermines it. So the firm runs unmonitored right up until a rare failure, then absorbs the entire cost in a single event it never saw coming. When that failure arrives, the cost does not fall evenly. It lands on the employee whose vigilance had been eroded by the system's reliability, and that employee is the one held responsible.

Supporting Evidence

Parasuraman and Riley named this problem in their 1997 review. They first described the pattern by which a reliable machine quietly relaxes the human watching it. Their review sits inside the older human-automation tradition of Sheridan, Bainbridge, and Endsley, who studied what happens to an operator pulled out of the active control loop. The evidence assembled here reaches well past a single literature survey. It draws on three controlled experiments that measured the effect directly, and on first-person testimony from people who have lived through the rare failure. These settings trace one idea across the laboratory and the cockpit: dependability and inattention tend to grow on the same stem.

The story begins where the term itself was coined. Parasuraman and Riley (1997) reviewed decades of work on how people use and misuse automation, and they isolated a specific failure mode. When a system proves consistently reliable, operators stop monitoring it for the rare lapse. Their review framed the puzzle, but it could not measure the puzzle on its own. Measurement came from the laboratory. Kaber and Endsley (1997) put operators in a dynamic control task. Pushing more of the work onto the machine pulled the human out of the loop, which weakened both awareness and the ability to step back in. Manzey, Reichenbach, and Onnasch (2012) sharpened the result. They raised the degree of automation in a process-control simulation and watched verification decline. The most-automated group checked the machine least, and when the aid finally gave a wrong answer, up to half of the participants followed it. Eleven of the eighteen who erred had already sampled the information that would have caught the mistake, which separates simple complacency from active bias.

The signature claim of this tension is that a near-perfect agentic system is harder to oversee than an unreliable one. Prinzel, DeVries, Freeman, and Mikulka (2001), in a NASA experiment, measured exactly that. Operators facing constant high reliability detected failures worse than operators facing variable reliability, a monitoring score of .70 against .84, while reporting lower effort. The reliability that lowered workload also dulled the watching.

The same shape appears outside the lab. Krikorian (2026), writing after his self-driving car crashed while driving perfectly, argued that a system which fails constantly keeps you sharp, while one that almost never fails is the dangerous one. Tudela (2026) described the human cost as functional sedation, a supervisor present in form but absent in substance. The thinner edge of the evidence deserves honest mention. The signature inversion rests on a single on-axis experiment. The mechanism is now shown rather than merely asserted, yet it would stand firmer on a second.

Business Tension

AI-side tensionHuman-side tension
Complacency induction versus Reliable operational capabilityFunctional sedation versus Effective command

3.4 · Results · Cluster 4 of 5

Automation-Induced Skill Atrophy

Automation-induced skill atrophy is the erosion of hands-on competence in tasks that once only humans could perform. The skill fades not by choice but through disuse, and the loss is progressive, accumulating quietly over time.

Business Impact

Once an agentic system takes over a task and runs it efficiently, the employees who used to do that work stop practicing it. That efficiency is exactly the return on the investment the company expected, but it is also what removes the daily practice that kept the human skill alive. The gain, though, shows up mostly at the individual desk and far less across the team. The more dependably the agentic system runs, the less reason anyone has to stay hands-on. To stay sharp, employees would have to periodically switch the automation off and do the work themselves. Deprived of that practice, the people hired to catch the agentic system's failures slowly lose the ability to do so. This leaves the company exposed, because on the day the agentic system breaks down, no one will be able to step in and fill the gap.

Supporting Evidence

Lisanne Bainbridge gave this problem its name in 1983, when she described the ironies of automation in industrial control rooms. Her point was simple and uncomfortable. The better a system runs itself, the less its human operator practices, so the person meant to take over in a crisis is the one least ready to. The evidence gathered here reaches from that founding essay to two controlled experiments, several large workforce surveys, and a set of practitioner accounts from software engineering. It spans more than four decades and several settings, from aviation-style control panels in the human-automation tradition of Endsley and Kaber, to present-day developers writing code beside AI assistants. Read together, these studies trace one mechanism across very different desks: reliable help quietly removes the practice that keeps people skilled.

The founding move is Bainbridge (1983). She watched operators supervise automated industrial processes, and she drew out an irony that organizes everything after it. Automation is meant to replace the unreliable human, yet it leaves that same human responsible for the rare emergency. Skills decay when they are not exercised, so a former expert who now only monitors slowly becomes a novice at the controls. Sustained vigilance is itself hard. She cited evidence that close monitoring breaks down after about half an hour. Her sharpest line is that the most successful automated systems, the ones that almost never need a human hand, may demand the heaviest investment in keeping that hand ready.

Kaber and Endsley (1997) turned the argument into measured data. Thirty subjects worked a dynamic control task across ten levels of automation, and the researchers timed how quickly each recovered when the system dropped out. Operators coming off high automation were slower to regain control, and they addressed fewer pending tasks. An intermediate level of automation preserved recovery skill and improved normal performance at the same time. That single experiment both demonstrates the decay the tension claims and points to the practice that prevents it. Two airline incidents from 1987 and 1989 ground the finding outside the laboratory.

The mechanism then crosses into knowledge work. Kosmyna and colleagues (2025) recorded the brains of essay writers over four months. The heaviest AI users showed the weakest fronto-parietal connectivity, the lowest sense of authorship, and almost no ability to quote their own sentences back. Reduced engagement persisted after the assistant was withdrawn, which is why the team called it cognitive debt. Their study also tested a fix. Writers who drafted first and consulted the model second kept their neural activation, a direct echo of the intermediate-automation remedy from aviation.

Different stakeholders fill out the picture. Among software developers, the Stack Overflow survey (2025) of roughly 29,000 respondents found that about 70 percent reported personal productivity gains from AI agents, yet only 17 percent saw better team collaboration. The efficiency lands at the individual desk far more than across the team. Khare (2026) supplies the human moment behind that statistic: faster shipping, then a stalled whiteboard. Walther (2025), writing from Wharton, frames the same arc as a four-stage slide toward learned technological helplessness, and proposes AI-free zones as a deliberate counterweight. At organizational scale, the Stanford AI Index (2025) records adoption climbing toward 78 percent, and Workday's survey (2025) of nearly 3,000 leaders finds 88 percent expecting agents to relieve workload. The skill-maintenance side has fewer measurements. Kaber and Endsley supply the one controlled win, while DeepMind's (2026) proposal to hand humans deliberate busywork remains a recommendation rather than a tested cost. That asymmetry is part of the honest reading here.

SourceDescriptionInsight
Kosmyna and colleagues (2025)Controlled EEG experiment with 54 participants writing essays over four months: heavy AI users showed the weakest neural engagement, the lowest sense of ownership, and an 83 percent failure to quote back their own writing, with the deficit persisting after the tool was withdrawn; a draft-first sequence preserved activation. Coins the term cognitive debt.The only direct neural measurement that AI reliance erodes the cognition that produced the work, and the only experimental test of a skill-maintenance intervention.
Kaber and Endsley (1997)Controlled experiment with ten levels of automation and 30 subjects: time-to-recovery and tasks handled under later manual control were worse after operators had functioned at high automation, while intermediate automation preserved recovery and improved normal performance. Two aviation incidents serve as field anchors.Directly measures the failure-recovery decay the tension claims and supplies the tested remedy in a single study.
Bainbridge (1983)Foundational essay naming the ironies of automation: physical skills deteriorate when they are not used; the most successful automated systems may need the greatest operator training; sustained monitoring is untenable past about 30 minutes; automatic control can camouflage failure until recovery is impossible.The definitional spine: the only source that formally names the trap where peak efficiency coincides with peak recovery incapacity.
Khare (2026)First-person practitioner account: work that took three hours now takes 45 minutes, followed by a whiteboard failure with no laptop and no AI, after months of outsourcing first-draft thinking to the tool.The clearest single moment in the corpus of skill decay surfacing precisely when the AI is absent.
Walther (2025)Wharton commentary naming a four-stage agency-decay cycle that ends in learned technological helplessness, with AI-free zones offered as the deliberate antidote.Frames skill atrophy as a discrete, progressive slide and supplies the progression half of the definition.
Ming (2026)Essay reporting a 72-participant forecasting experiment: most teams defaulted to the AI answer and gained nothing over the AI alone, while the small minority who demanded counter-arguments preserved their judgment.Names the efficiency-to-disengagement mechanism and measures it directly, at the same cognitive altitude as Kosmyna.
Stack Overflow (2025)Survey of roughly 29,000 developers: about 70 percent report individual productivity gains from AI agents, but only 17 percent report better team collaboration, a gap of more than 50 points.Measures the efficiency dividend and surfaces the individual-versus-team divergence that the firm-level story turns on.
DeepMind (2026)A framework proposing that AI systems deliberately assign humans tasks they could otherwise hand off, so people do not lose their skills; names the automation paradox and the intelligent-delegation problem.The strongest case that skill maintenance costs efficiency, but offered as a proposal rather than a measured cost.
Stanford AI Index (2025)Annual index reporting organizational AI adoption climbing toward 78 percent, alongside broad uptake among students and enterprises.Sizes the operational-efficiency pole at scale: the adoption that removes routine practice is widespread, not marginal.
Workday (2025)Global survey of 2,950 full-time decision-makers and implementation leaders: 88 percent expect AI agents to relieve workload, with strong twelve-month return expectations.Confirms the productivity dividend leaders are buying, the same dividend that quietly removes the daily practice keeping skills alive.

Business Tension

Companies must choose between optimizing agentic AI for maximum efficiency on work humans once performed themselves, or rotating those humans back through the task to hold off automation-induced skill atrophy so their competence doesn't drain away through disuse. Either way, the company pays: lean fully on the agentic system, and the day it fails, no one remembers how the work was done; preserve the skill, and you forfeit some of the savings the automation promised.

AI-side tensionHuman-side tension
Complacency induction versus Operational efficiencySkill atrophy versus Skill maintenance

3.4 · Results · Cluster 5 of 5

Cognitive Debt Accumulation

Cognitive debt accumulation is the lack of true, shared understanding of a product or system that a person or team experiences when an agentic system delivers finished, working artifacts for them at a much faster pace than anyone could produce or process themselves.

Business Impact

An agentic system can deliver results, presentations, codebases, applications, and designs, faster than any team working by hand. Yet precisely because it hands employees polished outputs, they skip the step-by-step reasoning that doing the work themselves would have required. So the team ships output without building a shared understanding of how it holds together. Deliberate reasoning, the kind that comes from working through a problem, develops most fully when a person actually does the work, and only partly when they receive its result. The agentic system removes that occasion. As a result, throughput rises while the team's understanding of its own systems may quietly thin, a latent cost the productivity numbers do not yet capture. This gap stays hidden until the work must be fixed or audited. By then, no one can clearly explain how it was built, and the cost comes due as heavier review, slower onboarding, eroded trust, longer hours, and the strain of supervising work no one fully understands. In the worst case, the knowledge cannot be reconstructed at all, and reviving the system costs as much as writing it from scratch.

Supporting Evidence

This problem has a named founder and a deep intellectual root, but only a narrow band of direct measurement. Margaret-Anne Storey carried the idea of cognitive debt into software engineering in 2026, and she built it on Peter Naur's 1985 argument that a working program is really a theory held in its builders' minds. The wider research here splits into three kinds. There is foundational theory from the software-engineering canon, in Naur and in Brooks before him. There is a layer of practitioner essays that describe the pattern from the field. And there are two controlled studies that measure something. One, by Cui and colleagues (2025), tracks how much faster developers ship with an AI assistant. The other, by Kosmyna and colleagues (2025), uses EEG to watch what happens to a person's own engagement when a model writes for them. The texture of the evidence is rich in concept and thin in direct proof.

The intellectual root runs deeper than AI. Naur (1985) defined a program as a theory living in its builders' minds, not as the text of its source code. His sharpest case was a compiler that a second team could not revive, despite complete code and documentation, because the people who held its theory had dispersed. The loss, in his account, was irreversible. Reviving the system could cost as much as writing it again. This is the substrate the whole pattern stands on, and it predates the agentic system entirely.

Storey (2026) translated that substrate into the AI era and carried the failure mode into software work. When an agent hands a developer finished, working code, the developer ships output without ever constructing the understanding that doing the work would have produced. Storey's classroom archetype, a student team paralyzed by week eight and unable to explain its own system, makes the abstraction concrete. Her account is argument rather than measurement, but it is the conceptual home of this problem and the source of its central image.

The productivity side is the best-measured part of the picture. Cui and colleagues (2025) report three pre-registered field experiments across 4,867 developers, with completed pull requests rising about 26 percent, and the steepest gains going to junior staff. Their quality proxies did not degrade: at one firm the approval rate even rose. This study confirms the velocity premise and, at the same time, pushes back on the cost thesis, because the one in-domain controlled trial measures output and finds no quality loss. The cost, if it exists, is in something Cui did not measure.

The one direct measurement of that cost sits outside software. Kosmyna and colleagues (2025) ran an EEG study with 54 participants over four months. Neural engagement scaled down as people offloaded more, highest for those writing unaided, then search users, lowest for model users. The model group could not quote text it had just generated and reported the weakest sense of ownership, and reduced connectivity persisted even after the tool was removed. Order mattered too. People who wrote first and used the tool later retained more than those who relied on it throughout. The reservation is that this is essay writing in education, not a development team, so the result has to travel across a domain gap to land on the claim it supports.

Two practitioner sources widen the setting. Osmani (2026) describes how the daily job moves from writing code to validating it, makes review the new bottleneck, and gathers the field's defect figures into a single account, prescribing a contract of never merging code the author cannot explain. Holbrook (2026) carries the same speed-versus-understanding tension into research work, where AI platforms flood organizations with high-volume, surface-level findings that look credible and quietly erode trust in research overall. Across theory, controlled study, and field report, every source points the same way, even as the heaviest cost claims still rest on a single out-of-domain measurement.

SourceDescriptionInsight
Cui and colleagues (2025)Three pre-registered field experiments with 4,867 developers: about 26 percent more completed pull requests with an AI assistant, with larger gains for juniors, and quality proxies showing no degradation.The one in-domain controlled study. It confirms the velocity premise and at the same time counter-weighs the cost thesis, because it measures output rather than comprehension and finds no quality loss.
Kosmyna and colleagues (2025)EEG experiment with 54 participants over four months: neural engagement scaled down as people offloaded more. Model users could not quote text they had just produced and felt the least ownership, with reduced connectivity persisting after the tool was removed.The only direct measurement of comprehension loss, and the sole primary behind every cost-side claim, but it studies essay writing in education rather than a software team.
Storey (2026)Applies cognitive debt to software engineering, over Naur's theory-building account, and supplies the archetype: a student team that hit a wall by weeks seven to eight and could no longer explain its own system. Argument, not measurement.The founding source of this problem and the origin of its central image, but it persuades by reasoning rather than data.
Naur (1985)Programming is theory-building: a program is a theory held in its builders' minds. When the people who hold that theory disperse, code and documentation alone cannot revive the system. Three qualitative cases.The axiomatic substrate for the skip-the-reasoning and irreversibility claims; the strongest theory source describes a loss closer to program death than to a repayable debt.
Osmani (2026)The job shifts from writing code to validating it, making review the bottleneck. Aggregates borrowed figures on AI-assisted code carrying more issues and a higher change-failure rate, and prescribes a contract of never merging code the author cannot explain.Supplies the defect and velocity figures the essays lack, though every number is cited from a primary rather than measured here.
Holbrook (2026)AI research platforms flood organizations with high-volume, surface-level findings that look credible but lack depth and context; prescribes deliberate counter-practices.Extends the speed-versus-understanding tension into research work and surfaces a systemic trust-erosion harm that the per-artifact claims do not capture.

Business Tension

Companies must choose between optimizing agentic AI for fast, finished deliverables and making teams do enough of the work themselves to avoid accumulating cognitive debt, so throughput doesn't outrun the understanding that keeps a system maintainable. Either way, the company pays: ship at the system's pace, and the team's grasp of its own work quietly erodes; slow down to build that shared understanding, and the productivity gain shrinks.

AI-side tensionHuman-side tension
Decontextualized automation versus Scalable consistencyCognitive passivity versus Deliberate reasoning

Taken together, these five descriptions are the catalog that answers RQ1: the recurring business tensions that arise from outsourcing human cognition to agentic AI. Each tension is defined, grounded in its own evidence base, and decomposed into an AI-side pole pair and a human-side pole pair. What the descriptions deliberately withhold is the verdict. Whether each tension is a trade-off to be decided or a paradox to be held is the work of Chapter IV, which runs the Dichotomy Probe to a verdict for each cluster and builds the Equilibrium Mapping the verdict calls for.