MBA Thesis • Chapter IV
This chapter answers the second research question for Outsourcing Human Cognition. The descriptive chapter named five recurring tensions that arise when firms hand cognition to agentic AI. This chapter classifies each one. The question for each tension is the same: is it a trade-off, where one pole is the defensible choice and the other is the mistake, or is it a paradox, where neither pole is defensible on its own and both must be held at once?
IV · Overview
For each of the five clusters, this chapter delivers the same sequence. It states the verdict. It poses the one-line Dichotomy Probe, the yes-or-no question that frames the tension. It runs a conclusion-first Probe Validation, in which both candidate answers are tested against the evidence, and both fail. It presents the Equilibrium Mapping, the two-armed balance of the AI system against the human factor, with the load-bearing pole on the left of each arm. It names the Smith and Lewis (2011) paradox category. It closes with a How Might We question that turns the verdict into a design brief and a short Transfer paragraph for managers.
The finding is uniform. All five tensions resolve to a paradox, not a dilemma. This is the analytical payoff that gives the paper its title: what managers reach for as either-or dilemmas are, on inspection, both-and paradoxes.
One structural reason explains why the verdict comes out this way so consistently. Frontier AI has what Karpathy (2024) calls jagged intelligence: a capability surface that is superhuman on some tasks and worse than a child on adjacent trivial ones, with no reliable way to predict in advance which task falls where. A dilemma needs one pole that is safe to choose. A paradox arises when neither pole is safe. On a jagged surface, "trust the system" fails because the hidden gaps bite, and "do not trust the system" fails because refusing the competent regime forfeits the larger gain. Neither pole survives, so paradox follows from the shape of the substrate, not from any weakness of the operator. The same logic predicts the exception. Where the jagged gaps are bounded by a narrow task surface and a verifiable output, tension can still emerge as a trade-off. That condition is met in exactly one place in the corpus, inside Automation Bias, which is why the corpus holds 61 paradox-typed pieces of evidence and 1 dilemma-typed one. The cluster-level verdict, the unit this chapter reports, is five paradoxes out of five.
The DEM category is named according to the cluster assigned by the Dynamic Equilibrium Model. The full exposition of that model, including all four categories and the paradox-versus-dilemma distinction, is given in Chapter II.
IV.1 · Analysis
Verdict: paradox. The load-bearing tension is miscalibrated trust against informed reliance.
Dichotomy Probe. When an agentic system advises on the decisions employees make and helps run their daily work, can they act on its recommendations without first checking how well it actually performs in their own work?
Neither answer survives contact with the evidence.
The first answer is yes: employees can act on what the system produces without first checking how well it performs in their own work. This does not hold, because a system's general capability is no guide to how it performs on a given task. In a field study of 758 BCG consultants, AI raised work quality by about 40 percent on tasks inside its competence, yet on a task placed just outside that zone, consultants using AI were 19 percentage points less likely to reach the right answer than colleagues working without it, who were correct 84.5 percent of the time (Dell'Acqua et al., 2023). The organizational picture matches. In Cortex's 2026 engineering benchmark of about 50 engineering leaders, 91 percent said AI had improved developer velocity and quality, but only 25 percent had data to support it, while the measured indicators moved the wrong way: incidents per pull request rose 23.5 percent, and change-failure rates climbed roughly 30 percent (Cortex, 2026). The employee cannot see which side of that line a given task falls on, and the system reads as equally fluent on both. So acting without checking is not defensible as a standing rule: it amounts to trusting a capability that may not be there.
The opposite answer is no: employees should withhold reliance until the system has proven how well it performs in their own work. This does not hold either, because that proof can never be fully reached. A jagged, shifting system cannot be certified in advance, because the next task may sit on the weak side of a frontier no one can see, so "prove it first" has no stable stopping point. A blanket demand for proof also pushes people into the opposite error, which is distrust and disuse. Dietvorst et al. (2015) found that after seeing an algorithm err, people abandoned it even when it still outperformed them, and Lee and See (2004) name this distrust as trust that falls short of the automation's capabilities. Withholding reliance until certainty is reached forfeits the gains the tool would have delivered, and it hardens into a reflexive refusal of a capable system. So a standing "verify before you ever rely" is not defensible either: it trades over-trust for an equally costly under-trust.
Neither end is defensible on its own. A standing yes over-trusts the system where it is weak. A standing no under-trusts it where it is quietly strong. Both fail because both forces are real: the system's capability is genuine, and the on-task reliability gap is genuine too. The tension cannot be settled by picking a side. It can only be managed by calibrating reliance on evidence of actual local performance, again and again. Because no fixed answer holds, the business tension is a paradox.
| AI SYSTEM | & | HUMAN FACTOR | ||
|---|---|---|---|---|
| 🅐 On-task reliability gap | 🅑 Capability-grounded transparency | 🅒 Miscalibrated trust | 🅓 Informed reliance |
DEM category: Learning. In Smith and Lewis (2011) terms, this is a learning tension: the system and its users have to keep adjusting trust as conditions change, because what the tool can do today is not a fixed quantity that can be settled once.
🅐 On-task reliability gap carries the load on the AI side. The cost is that a system's demonstrated ability outruns the reliability it shows once deployed. Klingbeil et al. (2024) and the Cortex (2026) benchmark both record this gap: output that looks competent enough to trust, yet is wrong often enough to hurt. Present conditions sharpen the cost, because agentic systems now produce smooth, well-phrased output at scale, and that polish is a habit learned in training, not proof the answer is right.
🅑 Capability-grounded transparency is the real counter-pole, not a wish. Okamura and Yamada (2020) built it into their experiment: an in-context signal of how reliable the system actually was, shown to operators while they worked. Every group that received the cue handed control back to a human at the right moments, while the no-cue control kept over-trusting. The paradox holds only because this gain is real: a clear reliability signal can pull trust back toward true performance.
🅒 Miscalibrated trust carries the load on the human side. This is confidence that does not track real reliability, and it persists. Klingbeil et al. (2024) found people kept over-trusting even when shown the AI could err. Bainbridge et al. (2011) and Glikson and Woolley (2020) explain why training and incentives do not erase it: presentation drives trust, so fluent output keeps earning confidence it has not earned on the merits.
🅓 Informed reliance is the human contribution no tool supplies on its own: reliance matched to actual performance. Okamura and Yamada (2020) measured it as appropriate manual takeover, the operator stepping in exactly when the system was weak. The clearest failure case is the mirror image: when the reliability signal was absent, the no-cue group never recovered, and the human judgment that would have caught the weak moments never engaged.
This paradox cannot be solved by choosing a pole, because neither pole is safe to choose. The useful question is therefore not "should people trust the system or not?" It is "how might we set things up so the two poles check each other rather than erode each other?" That shift turns a stuck choice into a design space a manager can act in.
How might we design agentic decision-making systems so that the system makes its own reliability visible and checks human miscalibration of trust, while informed human reliance checks a system that invites more trust than its on-task performance has earned?
For a manager, the lever is not a frontline operator trying harder to judge each output. It is whether the system shows its reliability while people use it. Okamura and Yamada (2020) demonstrate that an in-context reliability signal recalibrates trust where instruction does not. The practical move is to build capability-grounded transparency into the tool itself, so trust has something true to track. The tension is structural, so it is managed through how the system is built, not through telling people to trust more carefully.
IV.2 · Analysis
Verdict: paradox. The load-bearing human tension is cognitive passivity against operator vigilance, paired on the AI side with verification overload against decision accuracy. This cluster also holds the one dilemma-typed evidence in the corpus, in the bounded case where a single output can be checked against a clear rule. At the cluster level the verdict is paradox.
Dichotomy Probe. Can employees just follow an agentic system's usually correct recommendations without double-checking them, when it both recommends what to do and helps execute the work?
Neither answer survives contact with the evidence.
The first answer is yes: employees can act on the system's usually correct recommendations without checking them independently. This does not hold, because wrong recommendations occasionally occur, and they flow into actual decisions and leave damage behind. In the incentivized experiment by Klingbeil et al. (2024), people followed AI advice even when it contradicted both the available information and their own assessment, and their payoffs ran 22 percent lower than the no-advice control (Z = -2.962, p = 0.0028). The same failure appears in a safety-critical setting. Lyell et al. (2017) ran a controlled e-prescribing study with 120 medical students: when the decision support was wrong, omission errors rose 33.3 percent, and clinicians accepted false-positive alerts at high rates (most conditions p < .0001). The unchecked path also breaks at the system level. The Financial Times reported a 13-hour AWS outage after engineers let an agentic coding tool operate without intervention, and it autonomously deleted and recreated a live customer environment; a senior AWS employee called the failure small but entirely foreseeable (Rosner and Uddin, 2026). So following without an independent check is not defensible as a standing rule: it lets the system's rare error pass straight into a decision that harms the firm.
The opposite answer is no: employees must independently verify every output before acting on it. This does not hold either, because at scale no operator can audit every output, and the system gives no reliable signal of which output to check. This is the verification-overload pole, in which the same scale and speed that justify the investment overwhelm its validation. The model's own confidence cannot triage the work, because stated confidence does not track accuracy. Xiong et al. (2024) found that large language models are systematically overconfident, and even GPT-4's failure-prediction AUROC is merely 62.7 percent, barely above the 50 percent random-guess floor, so a confidently stated wrong answer reads the same as a right one. Full verification then cancels the benefit that justified adoption. Osmani (2025) reports that review, debugging, and testing do not compress just because code is generated faster, so the throughput gain is eaten by the checking, and the measured benefit of correct decision support that Lyell et al. (2017) recorded is forfeited if every output is treated as suspect. So checking outputs independently is not defensible either: it trades the rare-error risk for a verification burden that erases the throughput the tool was bought for.
Neither end is defensible on its own. A standing yes lets a rare wrong recommendation reach a real decision and cause harm. A standing no buries employees in checks that no one can complete and cancels the gains that justified the system. Both fail because both forces are real: at scale, the verification burden is genuine, and the human pull toward cognitive passivity at the verification gate is just as genuine. The tension can only be managed by deciding, case by case, which outputs receive independent scrutiny and which do not, as volume and stakes shift. Because no fixed answer holds, the business tension is a paradox.
| AI SYSTEM | & | HUMAN FACTOR | ||
|---|---|---|---|---|
| 🅐 Verification overload | 🅑 Decision accuracy | 🅒 Cognitive passivity | 🅓 Operator vigilance |
DEM category: Performing. In Smith and Lewis (2011) terms, a performing paradox is a clash between competing demands on output: here, the demand to ship more good work faster pulls against the demand to catch every error, and an organization has to satisfy both at once rather than pick one.
🅐 Verification overload is the load-bearing pole on the AI side. At scale, no human team can audit every output, and the model's own confidence cannot be trusted to flag the rare wrong answers that occur. Xiong et al. (2024) benchmarked five large language models and found them systematically overconfident, with average calibration error above 0.377. A confidently stated wrong answer looks just like a right one, so the burden of checking cannot be offloaded back onto the system itself. Shen (2025) shows a related failure: models supply answers when crucial information is missing rather than asking, which puts more unverified output in front of the employee.
🅑 Decision accuracy is the AI counter-pole, and the paradox only holds because this upside is real. When the decision support was correct, Lyell et al. (2017) measured prescribing errors falling by up to 46.6 percent, and the broader review by Lyell and Coiera (2016) confirms a reduction in decision errors. The Stack Overflow developer survey (2025) reports that roughly 70 percent of developers say agents contribute to higher personal productivity. These are evidenced gains, not marketing, which is why firms deploy these systems, and why simply switching verification on at full strength is not free.
🅒 Cognitive passivity is the liability on the human side, and it persists even under training, incentives, and correction. Klingbeil et al. (2024) paid participants for accuracy, yet they still over-relied on AI advice instead of their own judgment. Lyell et al. (2017) tested final-year medical students, a trained population, and found the same deference, with clinicians even overriding correct alerts 8.3 percent of the time. Task complexity and interruptions did not change the effect. The reflex to defer persists regardless of effort or expertise.
🅓 Operator vigilance is the human counter-pole and the irreplaceable human contribution. Independent verification is what catches the system's errors. Its absence is evident in failure cases. In the AWS outage, engineers let the agent resolve an issue without intervention, and the senior staff quoted called the failure small but entirely foreseeable (Rosner and Uddin, 2026). Lyell et al. (2017) name verification of alerts as the key safeguard against automation-bias errors. The human second look is the thing the system cannot supply for itself.
This paradox cannot be solved by either trusting the system fully or checking everything it produces, because both fail. The question therefore shifts from "trust or verify?" to "how might we hold both, so machine accuracy and human vigilance reinforce each other rather than cancel out?" That reframing opens a design space for where attention is spent.
How might we design agentic workflows so that machine accuracy offsets human passivity at the verification gate, while human vigilance offsets the verification load the system creates?
For a manager, this moves the issue off the employee's shoulders. The real concern is design-level: how the organization sets where attention is spent, what gets checked, and what gets waved through. Asking people to stay vigilant does not work when the model's own confidence cannot tell them which output is the dangerous one. The lever is structural. An organization manages this paradox by building verification into the system: routing scarce human checking to the high-stakes, judgment-heavy outputs where an error is costly, and letting structured tests absorb the rest.
IV.3 · Analysis
Verdict: paradox. The load-bearing tension is functional sedation against effective command.
Dichotomy Probe. When an agentic system reliably runs work that employees are meant to oversee, including high-stakes tasks where failures can cause real harm, can they keep monitoring it carefully even when its very reliability makes that monitoring feel unnecessary?
Neither answer survives contact with the evidence.
The first answer is yes: employees can relax their monitoring and give a passing sign-off once the system has proved reliable. This does not hold, because vigilance decays under proven reliability and the rare failure slips through. Banks et al. (2018) documented this pattern in an instrumented on-road study. Twelve experienced drivers drove a production Tesla Model S on Autopilot for about forty minutes on public roads, while four synchronized cameras captured the driver, the controls, the instrument cluster, and the road. A second analyst recoded a random sample to check the scheme, with agreement above 90 percent, and the analysis found clear signs of complacency and over-trust: the drivers were content to go fully hands-free, returning to the wheel only when prompted by system warnings. Krikorian (2026) shows where it can end: after months of fault-free trips on a familiar route, his Tesla Model X lost its bearings in a turn, and he grabbed the wheel too late to avoid a concrete wall. The laboratory tells the same story under controlled conditions. Manzey et al. (2012) found that the most automated operators verified the system the least, and when it gave an incorrect answer, up to half of them went along with it. Prinzel et al. (2001) measured it most directly: operators watching a steady, highly reliable system caught fewer failures than those watching an unreliable one, and they felt less burdened even as their watch was slipping. So a standing relaxation of monitoring is not defensible: it trades attention for reliability and then absorbs the whole cost of the one failure no one was watching for.
The opposite answer is no: employees must hold full active monitoring on every output for as long as the system runs. This does not hold either, because people cannot maintain effective vigilance over a highly reliable system that leaves their attention with nothing to do. Manzey et al. (2012) found that sampling of useful-but-optional parameters declined over time across all groups, not just the most-automated one, so attention drifted even where checking was free. Kaber and Endsley (1997) describe why: an operator relegated to passive monitor loses situation awareness, a role to which humans are ill-suited. Perpetual hand-monitoring also undermines the relief that justified the automation. Prinzel et al. (2001) measured that relief as lower workload under the reliable system (NASA-TLX 46.67 versus 57.05), and Manzey et al. (2012) found the most-automated group delivered the best routine performance at the lowest effort. Demanding constant vigilance forfeits those gains, and it still does not produce the watching it asks for. So a standing "watch everything, always" is not defensible either: it spends the effort the automation was meant to save and buys vigilance that human attention cannot actually sustain.
Neither end is defensible on its own. A standing yes lets reliability lull oversight into a sign-off that catches nothing. A standing no demands a constant watch that human attention cannot hold and that cancels the workload relief the system was bought to deliver. Both fail because both forces are real: the system's reliable operational capability is genuine, and that same reliability sedates the very attention meant to oversee it. The tension can only be managed by re-engaging human attention on the rare, high-stakes moments rather than asking for either blanket trust or unbroken watch. Because no fixed answer holds, the business tension is a paradox.
| AI SYSTEM | & | HUMAN FACTOR | ||
|---|---|---|---|---|
| 🅐 Complacency induction | 🅑 Reliable operational capability | 🅒 Functional sedation | 🅓 Effective command |
DEM category: Organizing. This means the tension is about how a firm arranges oversight and control: who watches the machine, how closely, and under what rules. It is not a one-time choice but a standing arrangement the organization has to keep tuning.
🅐 Complacency induction is the pole that carries the weight on the AI side. The system's own reliability is what pulls operators out of active supervision. Parasuraman and Riley (1997) first described this pattern, observing that when humans lean too heavily on automation, their oversight fades. Prinzel et al. (2001) later measured it directly: under constant high reliability, monitoring sensitivity dropped to .70, compared with .84 under variable reliability, a difference that was large and statistically significant (F(1,39) = 25.26, p < .0001). Present conditions sharpen this. Agentic systems now run longer chains of work with fewer visible checkpoints, so there is even less for a watching operator to catch, and the habit of watching fades faster.
🅑 Reliable operational capability is the counter-pole, and it is a real, evidenced gain, not a foil. The paradox only holds because this upside is genuine. The same NASA study found that the reliable system lowered operator workload (NASA-TLX 46.67 versus 57.05). Manzey et al. (2012) found that the most automated aid delivered the best routine performance and the lowest effort. Firms deploy these systems because the speed and throughput are worth having, which is why simply switching the automation off is not the answer.
🅒 Functional sedation is the load-bearing pole on the human side: oversight that is present on paper but absent in substance, the sign-off given without a real look. This pole persists even when firms try to train and incentivize it away. Manzey et al. (2012) found that operators sampled fewer of the optional checks over time across all groups. Kaber and Endsley (1997) describe the same drift in supervisory control: people kept in a passive monitoring role lose situational awareness and are slow to step in. Telling people to stay vigilant does not hold the pole back, because the system gives their attention nothing to do.
🅓 Effective command is the human counter-pole: real oversight, genuine situational awareness, and the capacity to intervene in time. This is the irreplaceable human contribution, and one concrete case shows what its absence costs. The driver who let his reliable car drive itself was no longer in effective command when it drove into a wall on a road he knew well (Krikorian, 2026). His name, not the maker's, landed on the insurance report. Effective command is exactly what the other three poles erode, and exactly what the firm still needs at the rare moment that counts.
This paradox cannot be solved by either resting on the system's reliability or demanding an unbroken watch, because both fail. The question shifts from "trust the machine or watch it constantly?" to "how might we hold both, so machine reliability and human judgment cover each other's blind spots?" That reframing opens a design space for oversight.
How might we design oversight in agentic systems so that the machine flags the moments that need a human, checking the natural decay of human vigilance, while human judgment counters the machine's quiet drift from reality?
For a manager, the costs and benefits arrive on different clocks: the efficiency gains are immediate and visible, while the costs are hidden and paid only when a failure eventually occurs. Even with the right training, incentives, and feedback, attention still drifts, because reliable systems leave vigilance with almost nothing to act on. The main lever lives at the design layer. Organizations can shape what shows up in front of employees, insert deliberate checkpoints that require real decisions, and create approval steps that only make sense after a thorough review. The remedy is design: building systems and oversight so that attention, challenge, and judgment are continuously required, instead of simply telling people to be more diligent.
IV.4 · Analysis
Verdict: paradox. The load-bearing tension is skill atrophy against skill maintenance.
Dichotomy Probe. If an agentic system consistently performs work under employee supervision, even in high-stakes roles where errors can cause harm, do those employees keep the practical skills needed to assume control immediately if the system fails?
Neither answer survives contact with the evidence.
The first answer is yes: let the system run the work and trust that supervising employees stay ready to take control the moment it fails. This does not hold, because hands-on skill decays when it is not used, and the takeover then arrives degraded. Haslbeck and Hoermann (2016) tested 126 randomly selected airline pilots on a manual raw-data precision approach. Pilots who relied on automation flew the manual approach less precisely, and recency of practice predicted manual flying skill more strongly than total flight experience or time since flight school. Onnasch, Wickens, Li, and Manzey (2014) found the same pattern across an integrated meta-analysis of 18 studies: as the degree of automation rises, the operator's situation awareness and return-to-manual performance degrade, and that loss surfaces exactly when the automation fails. Bainbridge (1983) names the underlying mechanism: an operator who has spent a long time only monitoring an automated process may now be an inexperienced one, because physical skills deteriorate without use, precisely when the failure demands someone more skilled, not less. So a standing yes is not defensible: it assumes a readiness that the supervision itself quietly erodes.
The opposite answer is no: keep employees doing the work by hand on a regular basis so their skills stay sharp. This does not hold either, because routine manual work forfeits the efficiency and workload relief that justified the automation in the first place. Onnasch et al. (2014) measured this trade-off across their 18 studies: a higher degree of automation improves routine task performance and progressively lowers operator workload. Reversing that, by keeping the operator hand-flying everything, gives the readiness back but sacrifices the productivity and workload dividend the system was bought to deliver, and it cannot scale to the full range of tasks the system now runs. The corpus does not contain a primary study that measures a skill-maintenance routine directly in terms of cost to throughput, so this cost rests on the readiness-versus-efficiency trade-off Onnasch measured rather than on a direct throughput measurement, and this limitation is acknowledged. So a standing no is not defensible either: it buys readiness at the expense of the efficiency that was the whole reason to automate.
Neither extreme holds up on its own. A standing yes lets human skills atrophy until the takeover fails. A standing no protects those skills but sacrifices the efficiency and workload reduction the automation was meant to bring. Both fail because both forces are real: reliable automation genuinely strips away the daily practice that sustains human expertise, and that same automation genuinely drives the throughput and reduced burden the firm invested in. The tension can only be managed by deciding how often to switch automation off and put people back in direct control, repeatedly trading a specific amount of efficiency for a specific amount of retained skill. Because no fixed answer holds, the business tension is a paradox.
| AI SYSTEM | & | HUMAN FACTOR | ||
|---|---|---|---|---|
| 🅐 Complacency induction | 🅑 Operational efficiency | 🅒 Skill atrophy | 🅓 Skill maintenance |
DEM category: Learning. In Smith and Lewis (2011) terms, a learning paradox is a tension between building new capability and sustaining the old. Here the firm builds capability into the machine, and the human capability that used to do the same work slowly erodes. The two pull against each other over time.
🅐 Complacency induction carries the tension on the AI side. Reliable automation removes the steady stream of feedback, vigilance, and active reasoning that kept the operator sharp. Onnasch et al. (2014) measured this across 18 studies: higher automation degraded operators' situation awareness and their ability to return to manual control. Haslbeck and Hoermann (2016) traced the same mechanism in the field, because the automation simply removed the practice that built the skill. Present conditions sharpen this cost, because agentic systems now run longer stretches of work with less human touch than the cockpit autopilots these studies examined.
🅑 Operational efficiency is the counter-pole on the AI side, and it is a real, measured gain, not a rhetorical concession. Onnasch et al. (2014) found a clear benefit: as automation increased, routine performance improved, and workload decreased. That dividend is why firms automate at all. The paradox only holds because this upside is genuine.
🅒 Skill atrophy carries the load on the human side: hands-on competence eroding through disuse. The clearest field measurement is Haslbeck and Hoermann's 126-pilot study (2016), where automation-reliant pilots flew a manual approach less precisely. Onnasch et al. (2014) add the return-to-manual degradation measured across their pooled studies. This pole persists under training and correction, because the decay tracks recent practice rather than total experience, so a well-trained operator who has stopped practicing still loses the skill.
🅓 Skill maintenance is the human counter-pole: the deliberate practice, periodic manual reversion, and recent hands-on work that keep the skill alive. Haslbeck and Hoermann (2016) showed that recency of practice predicted manual flying skill better than total experience, which means the remedy is practice, not seniority. Onnasch et al. (2014) point the same way, because lower degrees of automation preserved readiness. The human contribution that cannot be handed to the machine is the ability to take over at the moment of failure, the contribution Bainbridge (1983) warned would be missing exactly when it was needed most.
This paradox cannot be solved by either running the automation flat-out or forcing constant manual work, because both fail. The question shifts from "automate fully or practice constantly?" to "how might we hold both, so the efficiency the system earns pays for the practice that keeps people ready?" That reframing turns a stuck trade-off into a design brief.
How might we design agentic automation so that its operational efficiency pays for the skill maintenance that keeps deliberate practice sharp enough to catch the automation's failures?
For a manager, the upside and downside run on different clocks. The efficiency gains show up fast and loudly in dashboards and cost reports, while the loss of readiness creeps in slowly and silently, revealing itself only when a failure forces a manual takeover that staff are no longer fluent enough to handle. From the ground it looks like a frontline error, an operator fumbling a handoff. One level up it is a design choice: the organization has been running the system with no scheduled manual practice, steadily spending down its ability to recover. The real lever sits in how the work is designed. Readiness must be treated as a budget category to be financed, not assumed as a free side effect of having people in the room. That means scheduling manual practice on purpose, rotating operators back onto the controls before their skills fade, and building genuine reversions to human control into the workflow.
IV.5 · Analysis
Verdict: paradox. The load-bearing tension is decontextualized automation against scalable consistency.
Dichotomy Probe. When an agentic system generates most of the work a team ships, can that team keep the shared understanding it needs to learn and improve as it ships that work at scale?
Neither answer survives contact with the evidence.
The first answer is yes: let the system generate most of the work and assume the team's shared understanding survives the pace at which it ships. This does not hold, because understanding does not develop if the making process is skipped. In a randomized experiment built around a new Python library, developers who fully delegated coding to AI built about 17 percent less conceptual understanding of that library (Cohen's d = 0.738, p = 0.010), and the experiment found no statistically significant speed gain to offset that loss (Shen and Tamkin, 2026). The deficit then shows up in later work. When novices used AI without restrictions, they failed a subsequent maintenance task with the AI removed 77 percent of the time, compared with a 39 percent failure rate for novices who worked manually (Sankaranarayanan, 2026). The same pattern appears at the team scale. A student team shipping fast hit a wall in weeks 7 to 8 and could no longer make simple changes, because no one could explain why the system was built as it was (Storey, 2026). Naur (1985) names the underlying loss: a program is a theory held in the developers' minds, and once that theory is lost, the code and its documentation cannot revive it. So letting the system carry the work and trusting that understanding keeps up is not defensible as a standing rule.
The opposite answer is no: require that a human fully understand every artifact before it ships, so shared theory is never outrun. This does not hold either, because that rule forfeits the measured advantage that justified the tool and cannot keep pace at scale. The productivity gain is real and consistent. Across three field experiments with software developers, access to the AI coding assistant produced consistent productivity gains (Cui et al., 2025), and in a within-subjects study, students completed brownfield tasks about 35 percent faster with Copilot, p < 0.05 (Shihab et al., 2025). A standing demand for full prior human understanding of every artifact spends that speed back. At the scale at which an agentic system generates most of the work a team ships, full human comprehension of each artifact before it ships cannot be reached in time, so the rule throttles output to the speed of unaided review and discards the dividend the tool exists to deliver. So a standing "understand everything first" is not defensible either: it trades the cost of lost understanding for an equal and opposite loss of the speed that made the system worth adopting.
Neither end is defensible on its own. A standing yes ships at scale while the team's shared understanding quietly thins. A standing no protects that understanding but gives back the speed that justified the tool. Both fail because both forces are real: the system's scalable, consistent output is a genuine productivity dividend, and the deliberate reasoning that builds a team's shared theory of its own systems is genuinely load-bearing for the work that comes after. The tension can only be managed by deciding, case by case, where the team spends on understanding and where it spends on speed, and by deliberately rebuilding shared theory as the system grows. Because no fixed answer holds, the business tension is a paradox.
| AI SYSTEM | & | HUMAN FACTOR | ||
|---|---|---|---|---|
| 🅐 Decontextualized automation | 🅑 Scalable consistency | 🅒 Cognitive passivity | 🅓 Deliberate reasoning |
DEM category: Learning. Smith and Lewis (2011) use this category for tensions that turn on how knowledge and capability build over time. For this tension it means the conflict is about what a team learns and retains as it works, not just what it produces today.
🅐 Decontextualized automation carries the load on the AI side. The system produces structurally correct artifacts at machine speed, but it does not hand over the intent and conceptual structure behind them, which Naur (1985) called a program's theory. Shen and Tamkin (2026) measured this cost directly: full delegation produced working code and a weaker understanding of the library it used. Shihab et al. (2025) showed the same artifact-without-theory pattern, with the read-and-understand step displaced by prompt-and-view. Present conditions sharpen this. Agentic systems now generate most of what a team ships, so the volume of code that arrives without its theory keeps rising.
🅑 Scalable consistency is the counter-pole, and the paradox only holds because this gain is real. Cui et al. (2025) found consistent productivity gains from AI coding assistants across three field experiments. Shihab et al. (2025) found that students completing brownfield tasks were about 35 percent faster with Copilot and made more progress on solutions. This is the dividend that justifies deployment, and it is not a story the team can simply walk away from.
🅒 Cognitive passivity carries the load on the human side. People become editors who skim or do not read, and they stop reasoning. This pole persists even under correction and training. Shen and Tamkin (2026) found that delegation led to a comprehension deficit even though participants actively worked through the tasks. Sankaranarayanan (2026) showed how costly the passivity becomes downstream: under unrestricted AI, novices failed a later maintenance task with the AI removed 77 percent of the time. The passivity does not announce itself; it shows up when the support is gone.
🅓 Deliberate reasoning is the human counter-pole, the disciplined practice that keeps people actively reasoning, and it is the contribution AI cannot supply for the team. Sankaranarayanan (2026) built an Explanation Gate that made novices teach back the generated code before they could integrate it. That teach-back step cut the later failure rate from 77 percent to 39 percent, and it cost nothing in first-phase productivity. The clearest evidence for this practice rests on this single study, but it is in-domain and direct. Storey's (2026) classroom case anchors the same point from the failure side: the team that skipped this reasoning could not rebuild its shared theory when it finally needed to.
This paradox cannot be solved by either delegating freely or demanding full prior understanding of every artifact, because both fail. The question shifts from "ship fast or understand everything?" to "how might we hold both, so the system's speed and the team's understanding build each other rather than trade off?" That reframing opens a design space for how AI fits into the work.
How might we design agentic workflows so that the system steps in to help when human attention drifts into passivity, while deliberate human reasoning steps in when the system runs on autopilot without context?
For a manager, the workflows that make teams feel faster are the ones that quietly erode their ability to understand what they are doing. Short-term this looks like progress: more artifacts, less effort. Long-term it creates systems that no one can fully explain or safely change, so what felt like acceleration turns into drag. The implications are structural, not personal. If leaders treat AI mainly as a shortcut to thinking, they mortgage the organization's future ability to act with clarity. The shift is from cognitive delegation, where AI thinks and humans approve, to cognitive amplification, where humans think and AI helps express and test, so that using the tool preserves and extends human understanding rather than hollowing it out.
IV.6 · Synthesis
All five tensions came out the same way. Each is a paradox, not a dilemma. In each one, the Probe Validation showed both candidate answers failing against the evidence, and each Equilibrium Mapping balanced a genuine AI-side gain against a genuine human-side cost rather than naming a safe pole and a mistaken one.
The shape repeats because the cause is shared. Each cluster pulls on one scarce human faculty in two directions at once. Trust Miscalibration pulls on an accurate read of how reliable the system is. Automation Bias pulls on human attention at the moment of decision. Complacency pulls on sustained vigilance over time. Skill Atrophy pulls on practice, the hours of hands-on work that keep competence alive. Cognitive Debt pulls on shared understanding, the team's theory of its own work. In every case the same faculty is needed on both sides of the balance, and it cannot be given fully to the machine or fully to the human without a real loss on the other side.
The jagged-substrate argument explains why this holds across the set rather than by coincidence. Because the system's capability surface is uneven and its failures cannot be predicted in advance, neither blanket trust nor blanket distrust is defensible, so each tension settles into a both-and balance. The lone evidence-level dilemma inside Automation Bias is the exception that fits the rule: it sits in the bounded case where a single output can be checked against a clear standard, exactly where the substrate's gaps stop being jagged. At the cluster level, the unit this paper classifies, the verdict is five paradoxes out of five.
One common transfer thesis follows. In each cluster the real lever is not a frontline operator trying harder, but the design of the system and the work around it. The tensions are structural, so they are managed by building the cross-check into the workflow, where machine strengths cover human weaknesses and human strengths cover the machine's. Naming these tensions as paradoxes is therefore not a counsel of despair. It is the precondition for managing them well, and it sets up the design invitation the conclusion carries forward.