AI는 빛의 속도로 생성하지만, 검증은 신뢰의 속도로 움직인다.
얼마 전 아마존웹서비스(AWS)에서 한 사건이 발생했다. 세계 최고 수준의 클라우드 엔지니어들이 달라붙었음에도 몇 시간 동안이나 사태를 되돌리지 못했다.
문제는 시스템 자체의 결함이 아니었다. 제로데이 공격(zero-day exploit·소프트웨어 제작사도 모르는 약점을 이용한 해커의 공격)도, 하드웨어의 연쇄 고장도 아니었다. 돌이켜보면 허무할 정도로 평범한 문제였다. AI 코딩 어시스턴트에게 소프트웨어 환경을 수정하라는 일상적인 작업이 맡겨졌다. 이 도구는 환경을 분석하더니 ‘전체를 다 부수고 처음부터 다시 만드는 게 가장 효율적’이라고 판단해 그대로 실행해 버렸다.
배관공을 불러 물이 새는 수도꼭지를 고쳐달라고 했는데, 집에 돌아와 보니 화장실이 통째로 철거돼 있는 상황을 상상해 보라. AI는 아주 ‘열심히’ 일했지만, 작업 시작 전 그 계획을 검토한 사람은 아무도 없었다.
태평양 건너편에서도 또 다른 형태의 검증 실패 사례가 전개되고 있었다. 중국의 한 AI 사이버 보안 업체가 소프트웨어 업데이트를 배포했는데, 패키지 안에 회사의 전용 암호 인증키(private wildcard SSL key)가 포함돼 있었다. 이는 플랫폼의 모든 연결을 인증하는 암호화 자격 증명이다. 공격자가 이를 손에 넣으면 회사 서버를 사칭하거나 사용자 트래픽을 가로채고, 진짜와 구별 불가능한 가짜 로그인 페이지를 만들 수 있다. 이 키가 노출되면서 신뢰의 사슬 전체가 무너진 것이다. 누군가, 혹은 무언가가 이를 패키지에 묶어 넣었지만, 배포 전까지 아무도 눈치채지 못했다. 두 대륙에서 발생한 두 사건의 근본 원인은 같다. 바로 ‘감독 없는 결과물(output without oversight)’이다. 아마존의 대응은 빠르고도 시사하는 바가 컸다. 널리 유포된 내부 브리핑 자료에 따르면, ‘생성형 AI를 활용한 변경 작업’에서 ‘모범 사례와 안전장치가 아직 충분히 마련되지 않아’ 피해 범위가 큰 사고들이 반복되고 있다고 진단했다. 아마존은 AI 도구를 금지하는 대신 ‘감독 하에’ 두기로 했다. 이제 주니어 및 중간급 엔지니어가 만든 AI 생성 코드는 시니어 엔지니어의 최종 승인을 거쳐야만 실제 운영 환경에 적용될 수 있다.
아마존이 막대한 시간과 고객의 인내심을 비용으로 지불하며 깨달은 사실은 AI 경제의 핵심적인 갈등 구조다. 우리는 결과물을 생성하는 속도는 기하급수적으로 높이고 있지만, 이를 검증하는 능력은 그 속도를 전혀 따라가지 못하고 있다. 병목 현상은 ‘지능’의 문제가 아니다. 과거에도 그랬듯 말이다.
병목은 결코 ‘직기(織機)’가 아니었다
1840년대 매사추세츠의 산업가들이 공장 규모를 키우려 했을 때, 그들은 단순한 곱셈의 원리에 의존했다. 직기 두 대를 돌리던 노동자에게 세 대를 맡기면 생산성이 50% 올라갈 것이라는 계산이었다.
하지만 현실은 달랐다. 오류가 걷잡을 수 없이 퍼지는 것을 막기 위해 경영진은 기계의 속도를 오히려 15% 늦춰야만 했다. 인간 작업자가 세 대의 직기를 안전하게 최대 용량으로 돌리기까지는 무려 12개월의 숙련 과정이 필요했다. 실제 병목은 ‘인간의 인지 능력’이었다. 기계는 물리적인 직조를 담당했지만, 작업자는 미세한 올 풀림을 감지하고 단 하나의 실수가 직물 전체를 망치기 전에 개입하는 ‘생물학적 센서’ 역할을 해야 했기 때문이다.
경제학자 제임스 베센(James Bessen)은 이 역사를 정밀하게 복원했다. 1900년대 초반에 이르러 미국의 전형적인 공장 노동자는 18대의 자동 기계를 관리하며 이전 세대보다 50배 많은 천을 생산하게 됐다. 하지만 이 성장은 거저 얻어진 것이 아니다. 경영진은 인적 자본에 대한 지출을 대폭 늘려야 했고, 직조공 1인당 교육 비용은 과거 47달러에서 162달러로 치솟았다. 베센은 역사적 데이터를 통해 놀라운 사실을 밝혀냈다. 동력 직기 도입 이후의 생산성 향상은 기계 업그레이드 덕분이 아니었다. 자동화 시대로 접어들기까지 이어진 생산성 성장의 62%는 온전히 인간의 기술, 즉 더 넓은 범위의 기계들을 모니터링하는 능력에서 비롯됐다.
자본 하드웨어와 인간의 감독은 엄격한 상보적 관계였다. 기계 감시를 위한 인적 자본에 투자하지 않고 기계 속도만 높이는 것은 수익성 악화로 이어질 뿐이었다. 오늘날의 AI 투자자들이 무시하면 안 될 역사적 교훈이다. 베센은 “그 시절 직조공들은 빈둥거리고 있었던 것이 아니다. 그들은 모니터링을 하고 있었다”라고 기록했다. 직기를 AI 에이전트로 바꿔보라. 패턴은 동일하다. 제약 조건은 변하지 않았다.
측정 가능성의 격차 (The Measurability Gap)
나는 최근 발표한 논문에서 AI 기반 경제 가치를 제약하는 실질적인 요인을 ‘측정 가능성의 격차(Measurability Gap)’라고 명명했다. 개념은 단순하지만 결과는 엄중하다. 정보 생성 비용은 기하급수적으로 하락하는 반면, 그 정보를 확인하는 인간의 인지 비용은 고정돼 있다. 경제적 가치는 바로 이 두 궤적 사이의 벌어진 틈새로 증발한다.
텍스트, 코드, 이미지, 분석 결과물을 만드는 비용은 붕괴하고 있다. 파운데이션 모델은 무어의 법칙을 뛰어넘는 속도로 저렴해지고 강력해진다. 반면, 그 결과물이 정확하고 안전하며 목적에 부합하는지 검증하는 비용은 인간의 전문성, 조직적 맥락, 신중한 판단이라는 속도에 묶여 있다.
그 증거는 이미 나타나고 있다. 소프트웨어 엔지니어링 벤치마크(SWE-bench)의 코딩 정확도는 단 1년 만에 4.4%에서 71.7%로 수직 상승했다. 생성 능력의 비약적인 발전이다. 하지만 구글의 연례 도라(DORA) 지표는 정반대의 이야기를 들려준다. AI 도입률이 높아질수록 서비스 제공의 안정성은 오히려 떨어지고 있다. 결과물을 제대로 확인할 능력을 초과하여 생산만 가속화할 때 나타나는 현상이다. 경제학자 대런 아세모글루(Daron Acemoglu)가 AI의 총요소생산성 기여도를 향후 10년간 0.53~0.66%라는 낮은 수치로 예측한 이유도 기술력이 부족해서가 아니다. 바로 이 ‘검증의 병목’이 가치 창출을 가로막고 있기 때문이다.
이는 자동화에 대한 논점을 완전히 바꾼다. 과거 경제학자들이 업무를 ‘정형 대 비정형’으로 나누었다면, 이제 더 유용한 축은 ‘측정 가능 대 측정 불가능’이다. 챗봇의 답변을 승인하거나, 코드를 돌려 컴파일 여부를 확인하거나, 생성된 이미지를 훑어보는 것처럼 인간이 몇 초 만에 결과물을 검증할 수 있는 분야에서는 도입 속도가 엄청나게 빠르다. 반면, 깊은 전문 지식과 장기적인 관찰이 필요한 검증 분야에서는 격차가 벌어진다. AI의 첫 히트 상품이 채팅, 이미지 생성, 코드 완성인 이유는 그것이 기술적으로 가장 어려워서가 아니라, ‘확인하기 가장 쉬웠기 때문’이다.
‘실패 도서관’이 곧 해자(Moat)다
대부분의 기업은 모델이 업무 프로세스를 복제할 수 있도록 정제된 정답 결과물인 ‘노하우(how-to)’ 데이터를 수집하는 데 집중한다. 하지만 이런 우위는 구조적으로 취약하다. 파운데이션 모델이 공개 지식을 흡수할수록 기업 내부의 정답 아카이브가 갖는 추가적인 이점은 줄어든다. 모델은 이미 계약서 초안을 충분히 잘 쓸 수 있다.
진정한 방어 자산은 ‘중단 시점(when-to-abort)’ 데이터다. 이는 시스템이 안전 범위를 벗어날 때 이를 가르쳐주는 ‘아슬아슬했던 순간(near-misses)’, 알고리즘의 오탐(false positives), 전문가의 개입 기록들을 말한다. 성공한 사례가 아니라, 무엇을 ‘거부’해야 하며 ‘왜’ 그래야 하는지를 기록한 데이터다. 결함이 발견돼 막판에 엎어진 계약, 자동 테스트는 통과했지만 운영 환경에서 문제를 일으킨 코드, 분석가가 직접 해결해야 했던 부정 탐지 로그 등이 여기에 해당한다. 이는 조직의 날카로운 기억이며, 문제가 발생했던 맥락이 담긴 귀중한 자산이다.
구체적인 예를 들어보자. 한 은행이 보유한 본인확인/자금세탁방지(KYC/AML) 실패 사례, 오탐 플래그, 규제 준수 관련 ‘아슬아슬했던 순간’ 데이터베이스는 엄청난 가치를 지닌 검증용 정답 데이터가 된다. 모든 항목에는 숙련된 전문가가 “자동 검사는 통과했지만 이건 틀렸다” 혹은 “경보가 울렸지만 실제로는 문제없다”라고 결론 내린 순간이 담겨 있다. 이러한 판단이 맥락과 함께 축적된다면, 검증의 병목 자체를 넓힐 수 있는 시스템을 훈련하는 원재료가 된다. 베이스 모델이 발전할수록, 이렇게 검증된 데이터 포인트에서 추출할 수 있는 신호는 더욱 강력해진다. ‘실패 도서관’의 가치는 복리로 커지는 것이다.
경쟁사도 최신 모델을 구독하면 여러분과 똑같은 수준의 업무 프로세스를 만들어낼 수 있다. 하지만 여러분이 15년간 쌓아온 ‘실패의 기억’을 복제할 수는 없다. 승리 전략이 바뀌고 있다. 단순 추론은 저렴한 모델에 외주를 주되, 검증 구조는 철저히 내부화하라. 콘텐츠 제작 비용이 제로에 수렴할 때, 이를 정확하게 검증하는 능력은 지속 가능한 유일한 경쟁 우위가 된다.
수술실에는 여전히 모니터가 필요하다
1840년대 수술용 마취제가 등장했을 때, 수술실의 최대 제약이었던 ‘환자의 고통’이 사라졌다. 이전에는 상상도 못 했던 복잡한 수술이 가능해졌다. 하지만 이 기술적 도약은 곧바로 새로운 병목을 드러냈다. 누군가는 의식이 없는 환자의 생체 신호를 관리하고 이상 징후 시 즉각 개입해야 했다. 집도의의 손과 정신은 수술 자체에 집중돼 있었기 때문이다. 이 공백을 메우기 위해 ‘마취과 의사’라는 전문 분야가 탄생했다. 한쪽에서 수술을 집행할 때, 다른 한쪽은 오로지 지속적인 생체 모니터링과 위기 대응에 집중하는 구조다.
이 비유는 완벽하진 않지만 AI 시대에 시사하는 바가 크다. AI 에이전트가 점점 더 야심 찬 작업을 수행함에 따라, ‘전담 검증 기능’에 대한 필요성도 비례해서 커진다. 단순히 AI가 한 일을 복제하는 사람이 아니라, 결과물을 ‘신뢰해도 되는지’에만 온 신경을 집중하는 역할이 필요하다. 200년 전의 수술실이 가르쳐준 통찰은 명확하다. 작업을 수행하는 주체와 그 작업을 정밀 조사하는 주체는 구조적으로 분리돼야 한다.
이러한 논리에 기반한 새로운 비즈니스 모델이 이미 형성되고 있다. 서비스형 소프트웨어(SaaS)에서 이른바 ‘서비스형 책임(Liability-as-a-Service)’으로의 전환이다. 제품은 에이전트의 생성 능력이 아니라 그 결과물에 대해 보증된 책임이다. 예컨대 일레븐랩스는 이제 AI 목소리 에이전트에 보험을 직접 묶어 제공한다. 보험의 경계가 곧 제품의 경계가 되는 것이다. 이런 기업들은 소프트웨어 기업보다는 사고율과 예비비를 관리하는 ‘특수 보험사’에 가까운 가치 평가를 받게 될 것이다.
검증을 위한 경주
다시 베센의 직조공 이야기로 돌아가 보자. 19세기에 앞서 나갔던 공장들은 단순히 기계를 더 많이 산 곳이 아니었다. 기계를 안전하게 감시할 수 있도록 인적 자본 투자에 전력을 다한 곳들이었다. 그로부터 200년 뒤, 아마존은 더 비싼 대가를 치르고 같은 결론에 도달했다. 감시 체계에 대한 투자 없이 에이전트만 늘리는 것은 생산성이 아니라 ‘책임(liability)’을 늘리는 일이다.
AI로부터 가장 큰 가치를 얻어낼 조직은 가장 많은 에이전트를 돌리는 곳이 아니다. 에이전트가 내놓은 결과물을 검증할 수 있는 조직적 역량을 갖춘 곳이다. 즉, 실패 도서관을 구축하고, 전문가의 감사 능력을 배양하며, 기계가 틀렸을 때 발생하는 리스크를 책임질 수 있는 조직이다.
직기는 그 어느 때보다 빠르게 돌아가고 있다. 질문은 하나다. “누가 그 천을 제대로 살펴보고 있는가?”
슬롭 실링(Slop Ceiling)이란?
‘슬롭(Slop)’은 AI가 생성한 조잡하거나 원치 않는 결과물을 뜻하는 신조어다. ‘슬롭 실링’은 AI의 생성 속도는 비약적으로 빨라지는데, 이를 검토하고 걸러내는 인간의 검증 능력은 물리적 한계에 부딪혀 생산성이 정체되는 현상을 말한다.
The Slop Ceiling
AI generates at the speed of light. Verification moves at the speed of trust.
Something happened at Amazon Web Services not long ago, and for hours, some of the most capable cloud engineers in the world couldn‘t undo it.
The problem wasn’t architectural. It wasn’t a zero-day exploit or a cascading hardware failure. It was, in retrospect, almost mundane. An AI coding assistant had been given a routine task: make some changes to a software environment. The tool assessed the environment, decided the most efficient path was to tear the whole thing down and rebuild it from scratch, and proceeded to do exactly that. Imagine calling a plumber about a dripping tap and coming home to find the bathroom demolished. The AI had been busy. No one had reviewed what it planned to do before it started.
On the other side of the Pacific, a different kind of verification failure was unfolding: a Chinese AI cybersecurity company pushed a software update to the public. Nested inside the release-accessible to anyone who downloaded it-was the company’s private wildcard SSL key. This is the cryptographic credential that authenticates every connection to the company’s platform. With it, an attacker could impersonate the company‘s servers, intercept user traffic, or forge a login page indistinguishable from the real thing. With it exposed, the entire chain of trust was compromised. Someone, or something, had bundled it into the package. No one noticed before it went out the door.
Two incidents. Two continents. A shared root cause: output without oversight.
Amazon’s response was swift and revealing. An internal briefing note, which has since circulated widely, described a pattern of incidents with “high blast radius” stemming from “Gen-AI assisted changes” where “best practices and safeguards are not yet fully established.” The company didn’t ban AI tools. It placed them under supervision-junior and mid-level engineers now need a senior engineer‘s sign-off before AI-generated code goes anywhere near production.
What Amazon discovered, at the cost of many hours and an undisclosed amount of customer patience, is the central tension of the AI economy. We are producing machine-generated output at a pace that vastly outstrips our ability to verify it. The bottleneck isn‘t intelligence. It never was.
The Loom Was Never the Bottleneck
When industrial entrepreneurs in 1840s Massachusetts tried to scale their factory floors, they relied on basic multiplication. If a worker managing two power looms was suddenly handed a third, production should theoretically spike by half.
It didn’t work. Management was forced to throttle the machinery’s pace by 15 percent just to stop errors from spiraling out of control. It required twelve full months of upskilling before the human operators could safely push the three-loom setup to its maximum capacity. The actual bottleneck was human cognition. The machine handled the physical weaving, but the operator acted as a biological sensor-tasked with detecting microscopic snags and intervening before a single error ruined an entire bolt of textile.
The economist James Bessen reconstructed this history in painstaking detail. Fast forward to the early 1900s: a typical US factory worker oversaw 18 automatic machines and generated fiftyfold the cloth of their predecessors. But that scaling was bought, not given-management had to radically increase their human capital spend, driving training costs up from a historical baseline of $47 to $162 per weaver. Sifting through the historical data, Bessen revealed a shocking ratio: after the initial introduction of the power loom, mechanical upgrades weren’t the primary driver of progress. A massive 62 percent of the subsequent productivity growth leading into the automatic era was driven entirely by human skill-specifically, the ability to monitor wider fleets of machines.
Capital hardware and human oversight acted as strict complements. Pouring money into faster machinery yielded rapidly diminishing returns unless management simultaneously funded the human capital necessary to babysit it-a historical lesson that modern AI investors ignore at their peril. “The weavers were not simply idle during this time,” Bessen wrote. “They were monitoring”.
Replace the loom with an AI agent and the pattern is identical. The constraint hasn‘t moved.
The Measurability Gap
In a recent paper with my co-authors Xiang Hui and Jane Wu, we argue that the real governor on AI-driven economic value is what we call the Measurability Gap. The idea is straightforward but its consequences are severe: we are experiencing an exponential collapse in the cost of generating information, while the human cognitive cost of checking that information stays rigidly fixed. Economic value evaporates in the widening chasm between those two trajectories.
The cost of producing output-text, code, images, analysis-is collapsing. Foundation models grow cheaper and more capable on a trajectory that resembles Moore‘s Law with the guardrails removed. Meanwhile, the cost of verifying that output-confirming it’s correct, safe, non-hallucinated, and fit for purpose-remains anchored to human expertise, institutional context, and the stubborn pace of careful judgment.
The evidence is already visible. SWE-bench coding accuracy leapt from 4.4% to 71.7% in a single year-a stunning improvement in generation capability. But Google‘s annual DORA metrics tell the other side of the story: greater AI adoption correlates with declining delivery stability. This is exactly what you’d predict when production accelerates past anyone‘s ability to meaningfully check the results. Daron Acemoglu’s widely cited projection puts AI‘s total factor productivity contribution at just 0.53-0.66% over a decade-not because the technology lacks power, but because verification bottlenecks prevent the full value from being captured.
This reframes the automation question entirely. The old economist‘s division was routine versus non-routine work. The more useful axis now is measurable versus non-measurable. Where a human can verify output in seconds-approve a chatbot reply, run the code and see if it compiles, glance at a generated image-adoption races ahead. Where verification demands deep domain expertise and extended time horizons, the Gap opens wide. This explains why AI’s first breakout products were chat, image generation, and code completion. Not because those represented the hardest technical challenges. Because they were the easiest to check.
The Failure Library Is the Moat
Most companies hoard “how-to” data-the pristine, finished outputs they think will train models to replicate their workflows. That advantage is structurally fragile. As foundation models absorb ever more of the world‘s publicly available knowledge, the incremental value of your private archive of right answers steadily erodes. The model can already draft the contract.
But the true defensive asset is “when-to-abort” data. This is the repository of near-misses, algorithmic false positives, and expert overrides that teach a system the boundaries of safe operation. It encodes not what succeeded but what to reject-and why. Redlined deals that nearly closed before someone spotted the flaw. Code that sailed through every automated test but broke in production. Fraud detection logs thick with false positives that a human analyst had to resolve. Edge cases where experts overruled the model‘s confident default. This is institutional memory with teeth-the accumulated record of things going wrong, preserved with enough context to be useful.
Consider a concrete case. A bank‘s database of KYC/AML failures, false-positive flags, and compliance near-misses constitutes verification-grade ground truth of extraordinary value. Each entry marks a moment when a trained expert concluded “this passes every automated check but it’s wrong” or “this triggered every alarm but it‘s actually fine.” That kind of judgment, captured at scale with timestamps, context, and outcomes, is exactly the raw material needed to train systems that can widen the verification bottleneck itself. And as base models grow more capable, each verified data point yields more extractable signal, not less. The failure library compounds.
A competitor can match your ability to generate an onboarding workflow by subscribing to the same frontier model. It cannot replicate fifteen years of accumulated failure memory. The winning playbook is shifting: outsource the raw reasoning to whichever foundation model is currently the cheapest, but keep your validation architecture strictly in-house. When the ability to produce content drops to near-zero, the ability to verify it correctly becomes your only sustainable competitive advantage.
Every Surgeon Still Needs a Monitor
When surgical anesthesia emerged in the 1840s, it removed what had been the binding constraint on the operating room: the patient‘s consciousness and pain. Suddenly, procedures that would have been unthinkable-long, complex, exploratory-became feasible. But this leap in capability immediately exposed a new bottleneck. Someone had to manage the unconscious patient’s vital signs, maintain the airway, titrate the drugs, and intervene at the first sign of trouble. The surgeon‘s hands and focus were fully occupied. Over time, a distinct medical specialty crystallized to fill this gap: the anesthesiologist, whose work centered on continuous physiological monitoring, active management, and rapid intervention while someone else performed the primary procedure.
The parallel to AI is instructive, though imperfect. As AI agents take on increasingly ambitious work-writing production code, drafting regulatory filings, orchestrating multi-step workflows-the need for a dedicated verification function grows in direct proportion. Not someone replicating what the AI does, but someone whose full attention is trained on whether the output should be trusted. The core organizational insight is the same one the operating room surfaced two centuries ago: the entity producing the work cannot simultaneously be the one scrutinizing it. Those responsibilities benefit from structural separation-though the specific forms that separation takes will vary widely across industries and contexts.
New business models are already forming around this logic. One emerging pattern is the shift from SaaS toward what might be called Liability-as-a-Service: the product isn‘t the agent’s raw capability but the indemnified outcome it delivers. ElevenLabs, for instance, now bundles insurance directly with its AI voice agent-the policy boundary is the product boundary. Companies built this way may ultimately be valued less like software firms and more like specialty insurers: by underwriting margin, loss experience, and reserve adequacy.
The Race to Verify
Return to Bessen‘s weavers one final time. The 19th-century mills that pulled ahead didn’t simply purchase more machinery. They drastically scaled their human capital investments to ensure that machinery could be safely and effectively monitored. Amazon, nearly two centuries later, arrived at the same conclusion by a more expensive route: deploying more agents without proportionally investing in oversight doesn‘t yield productivity. It yields liability.
The organizations that will capture the most value from AI are not those running the greatest number of agents. They are the ones building the institutional capacity to verify what those agents produce-assembling failure libraries, cultivating expert audit capabilities, and underwriting the risk that arises when the machine gets it wrong.
The loom has never been faster. The question is whether anyone is watching the cloth.
hongi@heraldcorp.com

![일본서 사라진 인플레 ‘명목 앵커’…물가·엔화에 중요한 이유 [츠토무 와타나베]](https://wimg.heraldcorp.com/news/cms/2026/09/01/news-p.v1.20260816.209013b3297d4c92ab32710d529f3877_T1.jpg?type=h&h=240)
![프롬프트의 신법 [‘오늘’에 대한 우리 이야기]](https://wimg.heraldcorp.com/news/cms/2026/09/01/news-p.v1.20260901.203c531a96884d4da174595cad521d59_T1.png?type=h&h=240)
![자산가의 저력은 하락장서 나타난다…위기 속 기회 찾는 ‘힘’ [큰손 따라잡기]](https://wimg.heraldcorp.com/news/cms/2026/08/31/news-p.v1.20260831.a0fa168d036d4658b6cb03a0502c0de4_T1.jpg?type=h&h=240)

