<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Rishabh Arora</title>
  <subtitle>Essays and notes on cognition, systems, strategy, and the craft of thinking clearly — written to be verified, not admired.</subtitle>
  <link href="https://rishabharora-kk.github.io/blog/feed.xml" rel="self"/>
  <link href="https://rishabharora-kk.github.io/blog/"/>
  <updated>2026-09-14T09:00:00+05:30</updated>
  <id>https://rishabharora-kk.github.io/blog/</id>
  <author><name>Rishabh Arora</name></author>
  
  <entry>
    <title>Verification Is the Binding Constraint</title>
    <link href="https://rishabharora-kk.github.io/blog/verification-is-the-binding-constraint/"/>
    <updated>2026-09-14T09:00:00+05:30</updated>
    <id>https://rishabharora-kk.github.io/blog/verification-is-the-binding-constraint/</id>
    <summary>Skill does not increase without a cost for being wrong. Frontier labs pay a 6x spread inside one company for work that resists specification, which is a premium on verification rather than on production. One metric detects whether a working day generated any learning at all.</summary>
    <content type="html">&lt;p&gt;An operator’s value is bounded above by their capacity to verify their tools’ output. Below that bound, a more powerful tool amplifies errors at exactly the rate it amplifies successes, and the expected gain from an arbitrarily strong generator is zero.&lt;/p&gt;

&lt;p&gt;That bound is not obvious from inside, because the instrument people use to check it — the felt sense of moving fast — has been measured, and it is inverted.&lt;/p&gt;

&lt;nav class=&quot;toc&quot;&gt;
  &lt;h4 class=&quot;no_toc&quot; id=&quot;contents&quot;&gt;Contents&lt;/h4&gt;
&lt;ul id=&quot;markdown-toc&quot;&gt;
  &lt;li&gt;&lt;a href=&quot;#the-two-obvious-answers&quot; id=&quot;markdown-toc-the-two-obvious-answers&quot;&gt;The two obvious answers&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#what-the-frontier-prices&quot; id=&quot;markdown-toc-what-the-frontier-prices&quot;&gt;What the frontier prices&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#three-absences-and-why-the-innermost-one-is-invisible&quot; id=&quot;markdown-toc-three-absences-and-why-the-innermost-one-is-invisible&quot;&gt;Three absences, and why the innermost one is invisible&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#boredom-is-a-correct-measurement-pointed-at-the-wrong-variable&quot; id=&quot;markdown-toc-boredom-is-a-correct-measurement-pointed-at-the-wrong-variable&quot;&gt;Boredom is a correct measurement pointed at the wrong variable&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-waterline&quot; id=&quot;markdown-toc-the-waterline&quot;&gt;The waterline&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-metric&quot; id=&quot;markdown-toc-the-metric&quot;&gt;The metric&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#four-loops&quot; id=&quot;markdown-toc-four-loops&quot;&gt;Four loops&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#where-the-exponential-actually-lives&quot; id=&quot;markdown-toc-where-the-exponential-actually-lives&quot;&gt;Where the exponential actually lives&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#how-this-fails&quot; id=&quot;markdown-toc-how-this-fails&quot;&gt;How this fails&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#where-this-is-most-likely-wrong&quot; id=&quot;markdown-toc-where-this-is-most-likely-wrong&quot;&gt;Where this is most likely wrong&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#the-method-with-the-domain-stripped-out&quot; id=&quot;markdown-toc-the-method-with-the-domain-stripped-out&quot;&gt;The method, with the domain stripped out&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#one-sentence&quot; id=&quot;markdown-toc-one-sentence&quot;&gt;One sentence&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#sources&quot; id=&quot;markdown-toc-sources&quot;&gt;Sources&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;/nav&gt;

&lt;h2 id=&quot;the-two-obvious-answers&quot;&gt;The two obvious answers&lt;/h2&gt;

&lt;p&gt;There are two default responses to someone who has gone fluent with AI tools and stopped improving. This argument rejects both, so they are worth stating first, at their strongest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One:&lt;/strong&gt; the deficit is unassisted production, so the remedy is to stop leaning on the tools and practise by hand until production returns. &lt;strong&gt;Two:&lt;/strong&gt; pay at frontier labs tracks seniority and the difficulty of the labour, so the spread inside a company should be modest and the top of the band is mostly a reward for tenure.&lt;/p&gt;

&lt;p&gt;Both are checkable against public data. Both fail, and they fail in the same direction.&lt;/p&gt;

&lt;h2 id=&quot;what-the-frontier-prices&quot;&gt;What the frontier prices&lt;/h2&gt;

&lt;p&gt;Reported figures, from compensation aggregators and public filings. Treat them as order-of-magnitude rather than precise, and check the links at the end.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Anthropic’s reported band runs from roughly &lt;strong&gt;$199K&lt;/strong&gt; total compensation at the low end (Trust and Safety) to about &lt;strong&gt;$1.27M&lt;/strong&gt; at the high end (staff-level software engineer). Federal H-1B filings reported base salaries of &lt;strong&gt;$1.12M–$1.38M&lt;/strong&gt; for people titled &lt;em&gt;Member of Technical Staff&lt;/em&gt; — base cash, before equity.&lt;sup id=&quot;fnref:anthropic-comp&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-comp&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
  &lt;li&gt;OpenAI: Research Scientist total compensation reported around &lt;strong&gt;$763K (L4) to ~$1.39M (L5)&lt;/strong&gt;, median package around &lt;strong&gt;$1.25M&lt;/strong&gt;; base bands for Member of Technical Staff (Research) around &lt;strong&gt;$245K–$685K&lt;/strong&gt;. Senior retention grants have been reported in the &lt;strong&gt;$5M–$20M/yr&lt;/strong&gt; range.&lt;sup id=&quot;fnref:openai-comp&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:openai-comp&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
  &lt;li&gt;Both converge on the same title architecture — &lt;em&gt;Member of Technical Staff&lt;/em&gt;, &lt;em&gt;Research Engineer&lt;/em&gt;, &lt;em&gt;Research Scientist&lt;/em&gt; — with engineering and research on one ladder.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The top number is the least informative part. The signal is the &lt;strong&gt;ratio&lt;/strong&gt;: roughly 6x between the low and high band inside a single company. A spread that wide is not tracking seniority, and it is not tracking how hard the labor is. It tracks &lt;strong&gt;irreducibility&lt;/strong&gt; — how badly the work resists being specified in advance.&lt;/p&gt;

&lt;p&gt;The cheap end of the band is work that can be described before it is done. The expensive end is work where nobody yet knows what a correct answer looks like: alignment, interpretability, frontier training, reinforcement learning, large systems failing in novel ways. The premium is not on coding. It is on &lt;strong&gt;formulating and verifying&lt;/strong&gt; where there is no specification and no oracle.&lt;/p&gt;

&lt;h3 id=&quot;what-they-say-they-select-on&quot;&gt;What they say they select on&lt;/h3&gt;

&lt;p&gt;From Anthropic’s careers page:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“We care about what you can do, not where you learned to do it. About half our technical staff had no prior ML experience; about half have PhDs, but plenty of brilliant colleagues never went to college. &lt;strong&gt;If you’ve done interesting independent research, written a thoughtful blog post, or contributed to open source, put that at the top of your resume.&lt;/strong&gt;”&lt;sup id=&quot;fnref:anthropic-careers&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-careers&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Engineers here do lots of research, and researchers do lots of engineering… All our papers have engineers as authors, often as first author.”&lt;sup id=&quot;fnref:anthropic-careers:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-careers&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;From a Research Engineer / Alignment posting, under &lt;em&gt;Candidates need not have&lt;/em&gt;:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“100% of the skills needed to perform the job. Formal certifications or education credentials.”&lt;sup id=&quot;fnref:anthropic-re-alignment&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-re-alignment&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The &lt;em&gt;good fit&lt;/em&gt; list from that posting, in its order: significant software, ML, or research engineering experience; contributing to empirical AI research; familiarity with technical AI safety research; prefers fast-moving collaborative projects to extensive solo efforts; picks up slack even if it goes outside the job description; cares about the impacts of AI.&lt;sup id=&quot;fnref:anthropic-re-alignment:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-re-alignment&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;From OpenAI’s Research Engineer page:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“We’re looking for people with solid engineering skills (for example designing, implementing, and improving a massive-scale distributed machine learning system), writing bug-free machine learning code, and building the science behind the algorithms employed… engineers who are comfortable working in large distributed systems.”&lt;sup id=&quot;fnref:openai-re&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:openai-re&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3 id=&quot;decoding-the-postings&quot;&gt;Decoding the postings&lt;/h3&gt;

&lt;p&gt;Job-post prose, translated into the capability being purchased. This column is inference, not data — the quotes above are the data.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Stated requirement&lt;/th&gt;
      &lt;th&gt;Capability being bought&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;“design and run elegant, thorough experiments”&lt;/td&gt;
      &lt;td&gt;Converts a vague intuition into a test that could come out either way, and knows which outcome would refute it&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“understand why systems behave as they do”&lt;/td&gt;
      &lt;td&gt;Holds a causal model of an opaque system; debugs by hypothesis, not by mutation&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“writing bug-free ML code”&lt;/td&gt;
      &lt;td&gt;Has an internal error model — knows where bugs live before they manifest&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“massive-scale distributed systems”, “complex shared codebases”&lt;/td&gt;
      &lt;td&gt;Reasons about systems larger than working memory&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“papers have engineers as first author”&lt;/td&gt;
      &lt;td&gt;Can write. Prose quality is a readout of thought quality, and it is being selected on&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“picks up slack outside the job description”&lt;/td&gt;
      &lt;td&gt;Agency without a specification&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;“no credentials required; put independent research at the top”&lt;/td&gt;
      &lt;td&gt;Artifacts are the currency; credentials are an explicitly discounted fallback&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;strong&gt;Interview gates are gates, not the game.&lt;/strong&gt; Data structures and live coding are not a typing test and not a proxy for productivity. They are an unplugged verifier test: with the oracle removed, is there an internal model of the machine, and can a system be held in working memory and reasoned about? That is why hand-writing code is the measurement instrument — it is the only way to observe from outside whether the internal model exists. It is a floor, not a ceiling, and clearing a pass/fail gate cheaply is correct strategy rather than something to resent. It is also not where the compensation is.&lt;sup id=&quot;fnref:interview-gate&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:interview-gate&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The trajectory points at verification.&lt;/strong&gt; Anthropic has reported evaluating whether a model’s proposed next experimental step beats the human researcher’s, and states it reaches parity in a substantial fraction of real research sessions.&lt;sup id=&quot;fnref:anthropic-rsi&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-rsi&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; Generation of research &lt;em&gt;steps&lt;/em&gt; is being absorbed. What stays scarce inside the lab is what is scarce in any individual workflow: someone who can tell whether the machine’s output is right, and who decides what should be attempted at all.&lt;/p&gt;

&lt;h2 id=&quot;three-absences-and-why-the-innermost-one-is-invisible&quot;&gt;Three absences, and why the innermost one is invisible&lt;/h2&gt;

&lt;p&gt;Stagnation inside a high-activity workflow is not a defect in what is present. It is a missing mechanism — and three go missing together, nested, the outer ones cheap and visible, the inner ones load-bearing.&lt;/p&gt;

&lt;h3 id=&quot;recognition-is-not-production&quot;&gt;Recognition is not production&lt;/h3&gt;

&lt;p&gt;Reading correct code and producing it are distinct capacities. Recognition is cued retrieval against a stored trace; production is free recall plus construction. Recognition is systematically easier, and it is the standard source of overconfidence, because the phenomenology of &lt;em&gt;yes, that’s right&lt;/em&gt; is nearly identical to the phenomenology of &lt;em&gt;I could have written that&lt;/em&gt;. It is possible to read a language one cannot speak.&lt;/p&gt;

&lt;p&gt;The consequence is a working ceiling equal to the model’s ceiling, with the operator’s judgment contributing nothing.&lt;sup id=&quot;fnref:skill-atrophy&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:skill-atrophy&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt; In a market where everyone holds the same model, that is a commodity position with no growth term.&lt;/p&gt;

&lt;h3 id=&quot;the-instrument-is-broken-which-hides-the-first-absence&quot;&gt;The instrument is broken, which hides the first absence&lt;/h3&gt;

&lt;p&gt;The brain uses &lt;strong&gt;fluency as a heuristic cue for understanding&lt;/strong&gt;. Information that arrives smoothly is judged better-understood than information that arrives with effort, independent of actual comprehension. AI output is maximally fluent, so it maximally triggers the false cue. This is the machinery behind the illusion of explanatory depth — people confidently report understanding how a bicycle or a zipper works until asked to draw one.&lt;sup id=&quot;fnref:unsourced-fluency&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:unsourced-fluency&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;9&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;The industrial-scale version is the part worth internalizing precisely. In METR’s randomized controlled trial, experienced open-source developers working on real tasks in their own repositories were &lt;strong&gt;19% slower&lt;/strong&gt; with frontier AI tools. They had forecast a 24% speedup beforehand. Afterwards — having personally lived through the slowdown — they estimated they had been sped up by &lt;strong&gt;20%&lt;/strong&gt;. The belief survived direct experience of the opposite.&lt;sup id=&quot;fnref:metr&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:metr&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;10&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;The structure of that result matters more than the headline. It does not show that AI does not help; METR was explicit on that point, and different tasks, workflows, and models can and do differ. The transferable finding is narrower and less comfortable: &lt;strong&gt;the subjective sense of speed is not a measurement of speed, and it does not correct itself through experience.&lt;/strong&gt; The feeling is generated by the smoothness of the interaction, not by the throughput.&lt;sup id=&quot;fnref:perception-gap&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:perception-gap&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;11&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;The name for the surrounding state is &lt;strong&gt;metacognitive laziness&lt;/strong&gt;: fluent answers remove the difficulty signals that normally trigger self-monitoring, so verification stops being invoked at all. Anyone citing a feeling of efficiency as evidence that their method works is citing the one reading they are not licensed to trust.&lt;sup id=&quot;fnref:metacog-laziness&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:metacog-laziness&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;12&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;h3 id=&quot;there-is-no-loss-function&quot;&gt;There is no loss function&lt;/h3&gt;

&lt;p&gt;Learning requires an error signal with teeth. There are three sources of teeth: reality breaks something, other people judge it, or a prior commitment is contradicted. A tool-saturated solo workflow removes all three.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Reality.&lt;/strong&gt; The tool patches over gaps before they manifest as failures, so not knowing costs nothing locally. The bill is deferred and accrues as &lt;em&gt;cognitive debt&lt;/em&gt; — code produced faster than comprehension of it, until no one can say what the program does or how to change it safely.&lt;sup id=&quot;fnref:cognitive-debt&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:cognitive-debt&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;13&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Other people.&lt;/strong&gt; Nothing made is exposed to anyone with standing to call it bad.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Commitment.&lt;/strong&gt; Nothing is predicted before it is observed, so nothing can be surprising — and surprise is the only free source of labels.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result is not a small gradient. It is zero gradient. Years of activity with no shaping pressure produce exactly enthusiasm, breadth, fluency, and no movement. The system is behaving correctly given its inputs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skill is a function of consequence density.&lt;/strong&gt; There is no known mechanism by which competence increases in the absence of a cost for being wrong. A frictionless environment does not shape anyone, however sophisticated it looks from inside.&lt;/p&gt;

&lt;h2 id=&quot;boredom-is-a-correct-measurement-pointed-at-the-wrong-variable&quot;&gt;Boredom is a correct measurement pointed at the wrong variable&lt;/h2&gt;

&lt;p&gt;Two claims, both true, and the tension between them is where the leverage is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The boredom is accurate.&lt;/strong&gt; Tutorials and structured courses are genuinely low-value, and not because the learner is too advanced. It is structural: a tutorial is a guided path with no possibility of failure. It is engineered so the learner cannot be wrong, which means it cannot generate a single labeled error. It produces the sensation of learning — fluency, progress bars, working output — with near-zero information transfer. The boredom is detecting a real absence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The boredom misattributes the cause.&lt;/strong&gt; The conclusion usually drawn is &lt;em&gt;manual practice is redundant&lt;/em&gt;. The supported reading is &lt;em&gt;instruction-following is redundant&lt;/em&gt;. These feel identical from inside and are opposite in consequence. What made tutorials dead was the absence of falsification — and the absence of falsification is also what characterizes tool-saturated work. The common move is to flee one zero-feedback environment into a second zero-feedback environment with better aesthetics. That is why the stagnation follows.&lt;/p&gt;

&lt;p&gt;The underlying mechanism says which lever works. Dopaminergic signalling encodes &lt;strong&gt;reward prediction error&lt;/strong&gt; — the difference between expectation and outcome — not reward magnitude. Applied to learning, the operative quantity is &lt;em&gt;learning progress&lt;/em&gt;: the rate of change of competence, not its level. Intrinsically-motivated agents built on this principle abandon a task the moment the derivative flattens, regardless of the absolute value of the skill. This is the standard architecture of curiosity-driven learning and a well-documented account of human boredom.&lt;sup id=&quot;fnref:unsourced-rpe&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:unsourced-rpe&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;14&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;Which means the problem is not insufficient discipline, and resolve will not fix it. Resolve does not change a reward schedule; it spends finite willpower fighting one, a losing trade on any horizon longer than a few weeks. The intervention has to change the &lt;strong&gt;information geometry of the task&lt;/strong&gt; so that the derivative of &lt;em&gt;visible&lt;/em&gt; competence stays high.&lt;/p&gt;

&lt;p&gt;Three calibration points, because the reverse error is available and real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Difficulty must be desirable, not merely present.&lt;/strong&gt; Difficulty aids long-term retention while hurting short-term performance — generation over reading, testing over review, varied over blocked practice.&lt;sup id=&quot;fnref:bjork&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:bjork&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;15&lt;/a&gt;&lt;/sup&gt; But this reverses when working memory is already saturated: with high element-interactivity material and no scaffolding, added difficulty is &lt;em&gt;undesirable&lt;/em&gt; and produces load and nothing else.&lt;sup id=&quot;fnref:undesirable-difficulty&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:undesirable-difficulty&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;16&lt;/a&gt;&lt;/sup&gt; The target is the region where failure happens roughly a third of the time — enough to generate signal, not so much that the signal is noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI output is a worked example, which is the right scaffold at the wrong moment.&lt;/strong&gt; For a genuine novice, a complete worked solution beats problem-solving; that is the worked-example effect, and it means the intuition that AI-assisted work teaches &lt;em&gt;something&lt;/em&gt; is not baseless.&lt;sup id=&quot;fnref:worked-example&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:worked-example&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;17&lt;/a&gt;&lt;/sup&gt; But the effect reverses with expertise, and a &lt;em&gt;complete&lt;/em&gt; example is inferior to an &lt;strong&gt;incomplete&lt;/strong&gt; one. Examples with steps deliberately removed, which the learner must supply, produce better self-explanation and better transfer.&lt;sup id=&quot;fnref:worked-example:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:worked-example&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;17&lt;/a&gt;&lt;/sup&gt; That is the whole fix, stated technically: do not stop using worked examples — delete parts of them before reading, and supply the deletions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Surprise is the fuel.&lt;/strong&gt; The highest-density source of reward prediction error available to a learner is being wrong about something just committed to. Prediction-before-observation converts every observation into a labeled error signal. It manufactures consequence out of nothing, generates the surprise that &lt;em&gt;is&lt;/em&gt; the reward being chased, and measures calibration as a free side effect. It also restores the structure boredom needs: a game has an opponent and a scoreboard, and prediction supplies both.&lt;/p&gt;

&lt;p&gt;So the apparent dilemma between speed and depth is not a dilemma. The move is to &lt;strong&gt;insert a prediction between the request and the answer&lt;/strong&gt;. It costs seconds and converts the workflow from consumption into training.&lt;/p&gt;

&lt;h2 id=&quot;the-waterline&quot;&gt;The waterline&lt;/h2&gt;

&lt;p&gt;Capability commoditizes from the bottom of the abstraction stack upward. Each layer, once machine-cheap, stops paying.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;   owning the consequence        &amp;lt;- not automatable: requires a self with skin
   judgment under irreducible
     uncertainty                 &amp;lt;- the compensation premium lives here
   problem formulation           &amp;lt;- &quot;what is even the right question&quot;
   problem selection / taste     &amp;lt;- being actively measured, contested
-- waterline, moving up --------------------------
   architecture                  &amp;lt;- falling now
   implementation                &amp;lt;- mostly gone
   syntax / recall of APIs       &amp;lt;- gone
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The intuition that hand-writing code is a bad investment is &lt;strong&gt;correct about the bottom three layers&lt;/strong&gt;. It is a sound read of the trajectory. The error is concluding there is therefore nothing to invest in. There is; it sits above the waterline, and it is reachable by neither the tutorial route nor the pure-delegation route.&lt;/p&gt;

&lt;p&gt;This is also what happens as models improve. The scarce complement to a stronger generator is never a better generator — it is a better &lt;strong&gt;discriminator&lt;/strong&gt;. Generation is being supplied at collapsing marginal cost. Discrimination is not. The 6x compensation ratio, the job-post language about knowing &lt;em&gt;why&lt;/em&gt; systems behave as they do, and labs measuring research taste because it is the remaining bottleneck all point the same way.&lt;sup id=&quot;fnref:taste&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:taste&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;18&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;h3 id=&quot;why-the-discriminator-requires-having-been-a-generator&quot;&gt;Why the discriminator requires having been a generator&lt;/h3&gt;

&lt;p&gt;Here is the part that dissolves the tradeoff:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;A discriminator cannot be trained without generation. You have to have built the thing to know how it breaks.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;An error model — the intuition that says &lt;em&gt;this looks wrong&lt;/em&gt; before it can be articulated — is compiled out of one’s own failures, in one’s own hands, with one’s own bugs. It cannot be installed by reading, transferred by explanation, or borrowed from the model. It is the residue of having been wrong in a specific way often enough that the shape became recognizable. There is no shortcut, and the absence of a shortcut is good news: this capability does not commoditize.&lt;/p&gt;

&lt;p&gt;Which licenses the reclassification the whole argument turns on:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hand-writing code is not production. It is instrumentation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The point is not to obtain the code; at that the machine wins and will keep winning. The point is to &lt;strong&gt;calibrate a detector&lt;/strong&gt;, and the detector is the asset — the thing still worth something in five years, the thing the interview is probing, and the thing the top of the band is paying for. A calibration weight is not useful for building houses. That was never the objection to it.&lt;/p&gt;

&lt;h2 id=&quot;the-metric&quot;&gt;The metric&lt;/h2&gt;

&lt;p&gt;The highest-leverage intervention in a system is rarely the effort parameter. It is the information flow, and what gets measured.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;h3 id=&quot;maximize-labeled-errors-per-unit-time&quot;&gt;Maximize &lt;strong&gt;labeled errors per unit time.&lt;/strong&gt;&lt;/h3&gt;
  &lt;p&gt;A &lt;em&gt;labeled error&lt;/em&gt; is an occasion on which a specific expectation was committed to, and reality returned a verdict that could not be argued with.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The usual options sort themselves against it without further argument.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Activity&lt;/th&gt;
      &lt;th&gt;Labeled errors/hour&lt;/th&gt;
      &lt;th&gt;Why&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Watching a tutorial&lt;/td&gt;
      &lt;td&gt;~0&lt;/td&gt;
      &lt;td&gt;Engineered so the learner cannot be wrong&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Reading documentation&lt;/td&gt;
      &lt;td&gt;~0&lt;/td&gt;
      &lt;td&gt;No commitment, so no verdict&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AI writes it, skim, it runs&lt;/td&gt;
      &lt;td&gt;~0&lt;/td&gt;
      &lt;td&gt;The model absorbs the errors; yours never surface&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AI writes it, skim, it breaks, AI fixes it&lt;/td&gt;
      &lt;td&gt;~0&lt;/td&gt;
      &lt;td&gt;Still the model’s error signal, not the operator’s&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Predict output, run, diff&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;very high&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;Every line is a labeled trial, seconds per cycle&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Write from a blank file against a test suite&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;high&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;The failures are yours, the verdicts mechanical&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Predict the tool’s approach before asking, then diff&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;high&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;Free, continuous, fits an existing workflow&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Ship something people use&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;high, and unfakeable&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;Reality supplies labels nobody would think to request&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Publish reasoning where competent people can see it&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;high&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;Others find the errors one’s own model is blind to by construction&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Four properties carry the weight.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;It resists Goodhart.&lt;/strong&gt; Gaming it requires doing the work. Prediction accuracy on unseen code cannot be inflated without actually modelling the system. Compare lines written, hours studied, repositories starred — all trivially gameable, which is why they are the metrics people reach for.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It reframes failure as yield.&lt;/strong&gt; A day full of being wrong is a high-output day. That inverts the emotional sign of the obstacle, which is usually not laziness but aversion to visible incompetence.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It restores the derivative.&lt;/strong&gt; Labeled errors per hour is high-frequency, visible, and trends. It supplies a fast progress signal from a source that is real rather than illusory, which is why it does not require discipline to sustain.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;It is a defensible target for pride.&lt;/strong&gt; Not &lt;em&gt;I use AI well&lt;/em&gt; — a tool choice, and tools change underneath. Not &lt;em&gt;I know N languages&lt;/em&gt; — unfalsifiable. Instead: &lt;em&gt;my predictions about systems are calibrated, and I can tell when output is wrong.&lt;/em&gt; That claim is testable, and it is what the top of the band buys.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The corollary is the single most useful question to ask about any activity: &lt;strong&gt;what would tell me I was wrong, and how fast?&lt;/strong&gt; With no answer, the activity is entertainment. That is allowed. It should not be scored as growth.&lt;/p&gt;

&lt;h2 id=&quot;four-loops&quot;&gt;Four loops&lt;/h2&gt;

&lt;p&gt;Loops, not steps. Each has a verdict mechanism the practitioner cannot author, and each is instrumented so the derivative of competence is visible — the only thing that keeps anyone in the loop. None of them requires giving up the tool.&lt;/p&gt;

&lt;h3 id=&quot;1-interception--daily-nearly-free-the-keystone&quot;&gt;1. Interception — daily, nearly free, the keystone&lt;/h3&gt;

&lt;p&gt;Before asking the model for anything non-trivial, write down — in a file, not in your head, because unwritten predictions are retroactively edited — the approach you would take, the shape of the solution, the part you expect to be hard, and the specific place you expect a bug.&lt;/p&gt;

&lt;p&gt;Then ask. Then &lt;strong&gt;diff the prediction against the answer&lt;/strong&gt;, and log one line: where it matched, where it differed, and — separately, this distinction is the point — whether the difference was an error or merely a different valid choice.&lt;/p&gt;

&lt;p&gt;It costs under a minute. It converts an unlimited supply of fluent worked examples into an unlimited supply of labeled trials, rides a workflow already run dozens of times a day, and is a direct calibration measurement. Ten interceptions a day is roughly 3,000 labeled trials a year against a frontier model.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Instrumentation:&lt;/em&gt; weekly match rate. When it climbs, judgment is converging on the model’s. When it plateaus high, start predicting where the &lt;em&gt;model&lt;/em&gt; is wrong, and check. That is the graduation.&lt;/p&gt;

&lt;h3 id=&quot;2-deletion-practice--three-or-four-times-a-week-3045-minutes-bounded-on-purpose&quot;&gt;2. Deletion practice — three or four times a week, 30–45 minutes, bounded on purpose&lt;/h3&gt;

&lt;p&gt;Take something the tool built recently and that you understand in the recognition sense. Delete a functional piece. Reconstruct it from a blank file, against the existing tests, with the model closed.&lt;/p&gt;

&lt;p&gt;Start with deletions small enough to succeed about two times in three. Grow the deletion until that ratio holds again.&lt;/p&gt;

&lt;p&gt;This is the incomplete-worked-example effect: generation rather than review, which is the strongest desirable difficulty, scoped so working memory is not saturated, which is what keeps the difficulty desirable rather than merely painful.&lt;sup id=&quot;fnref:generation-effect&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:generation-effect&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;19&lt;/a&gt;&lt;/sup&gt; The verdict is mechanical and instant.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Instrumentation:&lt;/em&gt; deletion size reliably restorable, and time to restore. Both are numbers and both move weekly. This is the loop that puts a visible derivative on the missing capability — precisely what was absent when tutorials went flat.&lt;/p&gt;

&lt;p&gt;Hard rule: time-boxed. This is calibration, not production. It is supposed to end.&lt;/p&gt;

&lt;h3 id=&quot;3-the-artifact--continuous-and-the-actual-deliverable&quot;&gt;3. The artifact — continuous, and the actual deliverable&lt;/h3&gt;

&lt;p&gt;One real project, chosen on three criteria in priority order:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Free verdicts.&lt;/strong&gt; Success and failure are decided by something other than the author’s opinion — a number replicates or it does not, a test passes or it does not, a user returns or does not.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;High ceiling.&lt;/strong&gt; It absorbs ten years of increasing sophistication without changing identity.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Overlap with what is priced.&lt;/strong&gt; Empirical work on systems nobody fully understands.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two shapes satisfy all three unusually well. &lt;strong&gt;Reproduce a published mechanistic-interpretability result on a small model, then extend it by one question of your own&lt;/strong&gt; — verdicts are free, the extension is where judgment enters and becomes visible, the ceiling is unbounded, and it lands exactly on “interesting independent research.” Or &lt;strong&gt;build and publish an evaluation for a capability nobody has measured well&lt;/strong&gt; — eval-writing &lt;em&gt;is&lt;/em&gt; the discriminator skill turned into a product, the act of specifying what “correct” means where nobody has. It is in demand, legible, and achievable alone.&lt;/p&gt;

&lt;p&gt;Use the tool at full throttle inside this. That is not cheating; it is the working condition of the job. One rule: &lt;strong&gt;publish the reasoning, and own the claims.&lt;/strong&gt; A result whose grounds cannot be defended is a result that does not get stated.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Instrumentation:&lt;/em&gt; one public artifact per six weeks or so. Not a post &lt;em&gt;about&lt;/em&gt; something — a thing you did, what you found, what you expected, and where you were wrong.&lt;/p&gt;

&lt;h3 id=&quot;4-exposure--every-six-weeks-minimum&quot;&gt;4. Exposure — every six weeks, minimum&lt;/h3&gt;

&lt;p&gt;Put the artifact in front of people with standing to judge it, in a venue where they are free to call it wrong. Not an audience. A jury.&lt;/p&gt;

&lt;p&gt;It cannot be folded into the others, because an error model is blind exactly where it is broken, by construction. Only an external model finds those. This is also the entire mechanism of compounding, and the only loop that feels genuinely bad — which is the diagnostic that it is working, since it is the only one carrying real consequence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Interview gates&lt;/strong&gt;, separately: clear the data-structures and live-coding floor to a pass/fail standard, then stop. Treat it as a fixed cost paid efficiently, and do not let it become the thing done &lt;em&gt;instead of&lt;/em&gt; an artifact — that substitution is comfortable, legible, and fatal, because it looks like progress while producing nothing. Loops 1 and 2 clear most of that gate as a side effect, which is not a coincidence: the gate is a verifier test, and those loops build the verifier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Under scarcity:&lt;/strong&gt; if only one loop is sustainable, run Interception — nearly free, rides existing habits, and it repairs calibration, which is what makes everything else steerable. If two, add the Artifact, because without one there is nothing to compound. Exposure is the one that gets dropped and the one with the highest marginal value. That asymmetry is not an accident; it is aversion doing its job.&lt;/p&gt;

&lt;h2 id=&quot;where-the-exponential-actually-lives&quot;&gt;Where the exponential actually lives&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Skill is sigmoidal.&lt;/strong&gt; Within a domain, returns to practice follow a power law with a decelerating exponent: slow start, steep middle, plateau. Expecting exponential returns from this layer guarantees recurring disappointment — and that disappointment is itself a generator of boredom, because a sigmoid measured against an exponential expectation reads as failure of the method. It is not the method. It is the expectation applied to the wrong layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Artifacts compound superlinearly.&lt;/strong&gt; The artifact layer is a network process, and network processes exhibit preferential attachment: each visible, verifiable piece of work raises the probability of the next opportunity, and opportunities beget visibility. Reputation is power-law distributed. Skill is not.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Growth rate is not set by learning rate. It is set by artifact rate multiplied by the verifiability of each artifact.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three consequences.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Breadth without depth cannot compound.&lt;/strong&gt; Transfer happens from deep structure, never from surface familiarity. Ten shallow domains transfer nothing to an eleventh; one deep domain transfers to most. Breadth is not wasted, but it is not capital until one strand goes deep enough to generate transferable structure.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;An unpublished artifact has a multiplier of zero.&lt;/strong&gt; The quality of the work does not matter. Unexposed work generates no network effect and no external error-correction. Which is why Exposure is non-negotiable despite being the loop that gets skipped.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The compounding layer and the hiring layer are the same layer.&lt;/strong&gt; Independent research, a thoughtful blog post, open-source contributions, at the top of the resume, credentials explicitly discounted — the thing that compounds &lt;em&gt;is&lt;/em&gt; the thing that gets people hired. That convergence is rare and worth exploiting.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One concrete route: Anthropic runs a &lt;strong&gt;Fellows Program&lt;/strong&gt;, with tracks including AI Safety and Security, ML Systems and Reinforcement Learning, and Economics and Policy — a structured entry that selects on demonstrated capability rather than credentials. They state they do not run internships, and permit re-application after 12 months.&lt;sup id=&quot;fnref:anthropic-fellows&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:anthropic-fellows&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;20&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;On distance, without encouragement: the gap between no shipped work and no published reasoning, on one side, and a frontier-lab research role on the other, is large, and it is measured in years of these loops rather than months. But the &lt;em&gt;shape&lt;/em&gt; of the gap is favorable, and shape matters more than size. It is not credentialist, not political, and not gated behind anything that cannot be started today. It is a production gap, and production gaps close monotonically under a working loop. What is missing is one habit — the testing habit — and a habit is installable in a way a talent is not.&lt;/p&gt;

&lt;h2 id=&quot;how-this-fails&quot;&gt;How this fails&lt;/h2&gt;

&lt;p&gt;Named failure modes are catchable from inside while they happen. Unnamed ones are visible only in retrospect, which is too late.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Meta-procrastination.&lt;/strong&gt; Reading a sharp analysis of a problem delivers the same fluency reward as solving it, at a fraction of the cost. This is the highest-probability failure by a wide margin, and the recursion is exact: an essay diagnosing an offloading habit is a perfect vehicle for it. &lt;em&gt;Detector:&lt;/em&gt; 72 hours pass with nothing logged. &lt;em&gt;Countermeasure:&lt;/em&gt; one interception today, before this feels complete. Feeling finished is the symptom.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Converting loops into a curriculum.&lt;/strong&gt; Build a schedule, a tracker, a plan; feel the tutorial boredom arrive on cue; conclude the method failed. It did not — it was converted into the thing that fails. &lt;em&gt;Detector:&lt;/em&gt; optimizing the system instead of running it. Building tooling for one’s practice is procrastination in a lab coat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Manual practice later, after this project.”&lt;/strong&gt; It does not happen. Consequence-free environments do not spontaneously acquire consequence, and a project already running has no slot where difficulty naturally enters. The interception goes inside current work now, or not at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Undesirable difficulty.&lt;/strong&gt; Overcorrecting — no tools at all, a project far beyond scope, a hard language from scratch. Working memory saturates, nothing is learned, and the result is a legitimate-seeming failure that generalizes into &lt;em&gt;I tried the manual thing, it doesn’t work for me.&lt;/em&gt; Stay in the two-of-three success band.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity defense.&lt;/strong&gt; Pride attached to a tool choice generates motivated reasoning against all of this, and it arrives disguised as sophistication: &lt;em&gt;this overweights manual skill, the models will be better in a year, verification will be automated too.&lt;/em&gt; Some of that is partly true, which is what makes it effective. &lt;em&gt;Detector:&lt;/em&gt; arguing with the argument instead of running a loop. The argument may even be correct, and it is still, functionally, avoidance. A prediction match rate does not care about anyone’s position.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Goodharting the metric.&lt;/strong&gt; Logging easy predictions and scoring high. &lt;em&gt;Countermeasure:&lt;/em&gt; the metric is labeled errors, not correct predictions. A week with a rising match rate on trivial predictions is a failing week. Predict the things you expect to get wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fleeing at the plateau.&lt;/strong&gt; Deletion-practice numbers climb fast and then flatten, and the boredom instrument fires on schedule. What it measures is the derivative of &lt;em&gt;visible&lt;/em&gt; competence, and the fix is to change the difficulty, not the domain. Domain-switching at the plateau is the exact behavior that produces a decade of breadth. It feels like curiosity. It is the plateau.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Publishing polish instead of process.&lt;/strong&gt; Wanting the artifact to look impressive means hiding the errors, which removes the only part with value — and specifically the part the labs are reading for. A strong research submission reads like a methods section: precise, honest, explicit about what the work does and does not show.&lt;/p&gt;

&lt;h2 id=&quot;where-this-is-most-likely-wrong&quot;&gt;Where this is most likely wrong&lt;/h2&gt;

&lt;p&gt;First, the three framings this argument had to reject. Each is more plausible than what replaced it, which is why they survive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool dependence is not the variable.&lt;/strong&gt; Inability to write code from a blank file reads like the bottleneck. It is an observable of the bottleneck. Hand someone unassisted production tomorrow, change nothing else, and they still stagnate — they still have no mechanism that reports when they are wrong. People leaning on these tools harder than average do compound fast, when they check. The variable separating the cases is verification, not autonomy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pay tracking seniority.&lt;/strong&gt; It tracks irreducibility instead, and the 6x internal spread is what gives that away. Reading the top of the band — where attention naturally goes — hides it completely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The growth curve, on the wrong layer.&lt;/strong&gt; Skill acquisition read as exponential, with the shortfall read as failure of method. Skill is sigmoidal. The exponential is real and lives one layer up, in artifacts, which is a network process and not a practice curve.&lt;/p&gt;

&lt;p&gt;Then the things this does not settle. First, the compensation figures are aggregator- and filing-derived, not audited; the 6x ratio is robust to a fair amount of noise in the individual numbers, but I have not verified any single figure against a payroll record. Second, METR’s result is one trial, on experienced open-source developers, in repositories they already knew — the specific 19% does not transfer to other populations, and METR says so; what I am claiming transfers is the &lt;em&gt;direction of the perception gap&lt;/em&gt;, which is a weaker and better-supported claim. Third, the strongest counterargument is the one in the identity-defense section: if verification itself is automated to a high standard, the scarce complement moves again, and the discriminator loses its premium. I think that is the right thing to watch, and I do not think the argument survives unchanged if it happens. What would refute this piece is a demonstration that operators who cannot verify output nonetheless capture the gains from stronger tools.&lt;/p&gt;

&lt;p&gt;One disclosure about standing, because it changes how the second half should be read. This is an analysis of a mechanism, not a report from someone who has run these four loops to completion and measured the result. The loops are derived from the cited literature and from the structure of the argument. The empirical claims above carry sources; the loops carry only reasoning, and they should be read as a proposed protocol rather than a tested one.&lt;/p&gt;

&lt;h2 id=&quot;the-method-with-the-domain-stripped-out&quot;&gt;The method, with the domain stripped out&lt;/h2&gt;

&lt;p&gt;The moves, in reusable form.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Demand the loss function.&lt;/strong&gt; What generates verdicts here, who authors them, and how fast do they arrive? No loss function means no learning, however much activity there is. Start here.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Distrust the stated problem.&lt;/strong&gt; A self-reported bottleneck has already passed through the reporter’s model, which is the thing under examination. The presenting complaint is evidence about the complainant’s model, not about the system. Ask what would have to be true for the stated problem to be the real one.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Look for the missing mechanism, not the present defect.&lt;/strong&gt; Absence is systematically harder to see than presence, which is how it survives years of introspection.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Find the self-refuting evidence.&lt;/strong&gt; When an argument rests on a subjective measurement, go and find out whether that measurement is calibrated. This move is available far more often than it is used.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Read the spread, not the level.&lt;/strong&gt; A 6x ratio inside one organization is a stronger signal than any absolute number, because it reveals what that organization treats as substitutable. Levels are noisy; ratios are structural.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Project the commoditization frontier.&lt;/strong&gt; For any capability: which layer of the stack is this, where is the machine waterline, and which way is it moving? That converts &lt;em&gt;is X worth learning&lt;/em&gt; — unanswerable — into a positional question with a defensible answer.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Find the scarce complement.&lt;/strong&gt; When an input’s cost collapses, value migrates to the new binding constraint. Do not ask what is valuable; ask what the newly-abundant thing now requires more of.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reclassify before you optimize.&lt;/strong&gt; A forced choice usually encodes a mislabeled category. Most dilemmas are classification errors.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Collapse to one metric, then stress-test it against Goodhart.&lt;/strong&gt; If gaming the metric is easier than satisfying it, the metric is worse than nothing.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Solve for sustainability at the mechanism level.&lt;/strong&gt; Any prescription requiring someone to want what they do not want will fail. Knowing &lt;em&gt;why&lt;/em&gt; they do not want it is the design specification for one that does not need willpower.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Pre-label the failure modes&lt;/strong&gt;, including your own motivated reasoning, which is the one you will not catch otherwise.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;one-sentence&quot;&gt;One sentence&lt;/h2&gt;

&lt;p&gt;Optimizing unlabeled output rate produces fluency and no movement. Invert it — optimize labeled error rate, keep every tool, relocate pride from the tool to the detector, and publish the residue.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;

&lt;p&gt;Compensation and hiring data:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.levels.fyi/companies/anthropic/salaries&quot;&gt;Anthropic Salaries — Levels.fyi&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.levels.fyi/companies/anthropic/salaries/software-engineer&quot;&gt;Anthropic Software Engineer Salary — Levels.fyi&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.metaintro.com/blog/anthropic-1-3-million-ai-roles-salary-2026&quot;&gt;Anthropic Is Paying Up to $1.3 Million for Some AI Roles in 2026 — Metaintro&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://ctaio.dev/en/salary/anthropic-salary/&quot;&gt;Anthropic Salary (2026): Engineer &amp;amp; Researcher Compensation Bands — CTAIO&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.levels.fyi/companies/openai/salaries&quot;&gt;OpenAI Salaries — Levels.fyi&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://jobsbyculture.com/blog/openai-compensation-2026&quot;&gt;OpenAI Salary 2026: $249K–$1.28M TC, PPU Equity — JobsByCulture&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://ctaio.dev/en/salary/openai-salary/&quot;&gt;OpenAI Salary (2026): L4–L7 Compensation, PPUs &amp;amp; Signing Bonuses — CTAIO&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.entrepreneur.com/business-news/how-much-openai-employees-make-salaries-685000&quot;&gt;How Much OpenAI Employees Make — Entrepreneur&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Primary hiring-criteria sources:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.anthropic.com/careers&quot;&gt;Careers — Anthropic&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://job-boards.greenhouse.io/anthropic/jobs/4631822008&quot;&gt;Research Engineer / Scientist, Alignment — Anthropic (Greenhouse)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://job-boards.greenhouse.io/anthropic/jobs/4610158008&quot;&gt;Research Engineer / Scientist, Alignment, London — Anthropic (Greenhouse)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://job-boards.greenhouse.io/anthropic/jobs/5023394008&quot;&gt;Anthropic Fellows Program — Anthropic (Greenhouse)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.anthropic.com/careers/jobs&quot;&gt;Jobs — Anthropic&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://openai.com/careers/research-engineer&quot;&gt;Research Engineer — OpenAI&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.sundeepteki.org/advice/anthropic-research-engineer-interview-2026&quot;&gt;Anthropic Research Engineer Interview 2026 — Sundeep Teki&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.anthropic.com/institute/recursive-self-improvement&quot;&gt;When AI builds itself — Anthropic&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.wenbo.io/blog/taste-scaling/&quot;&gt;Taste Is the Next Capability AI Will Crack — Wenbo Pan&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Perception–reality gap and cognitive offloading:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://metr.org/blog/2026-02-24-uplift-update/&quot;&gt;METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://letsdatascience.com/blog/developers-thought-ai-made-them-faster-the-data-said-otherwise&quot;&gt;AI Coding Tools Made Developers 19% Slower: METR Study&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/pdf/2601.02410&quot;&gt;The Vibe-Check Protocol: Quantifying Cognitive Offloading in AI Programming (arXiv)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/pdf/2606.20598&quot;&gt;Using Biometrics to Understand AI-Assisted Coding Performance and its Perception (arXiv)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S2451958826001764&quot;&gt;AI-overdependence and human cognitive decline — ScienceDirect&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/html/2602.20206v2&quot;&gt;Mitigating “Epistemic Debt” in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts (arXiv)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://tianpan.co/blog/2026-04-19-skill-atrophy-ai-augmented-engineering&quot;&gt;The Skill Atrophy Trap — TianPan.co&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/pdf/2603.10025&quot;&gt;A Review of the Negative Effects of Digital Technology on Cognition (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Learning mechanism:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.columbia.edu/cu/psychology/metcalfe/PDFs/Metcalfe-BjorkVolSubmitFeb14Final.pdf&quot;&gt;Metcalfe &amp;amp; Bjork — Desirable Difficulties and Studying in the Region of Proximal Learning (PDF)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.01623/xml&quot;&gt;Desirable difficulties, worked examples, and completion problems — Frontiers in Psychology&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6099118/&quot;&gt;Undesirable Difficulty Effects in the Learning of High-Element-Interactivity Materials — PMC&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.structural-learning.com/post/generation-effect-active-learning&quot;&gt;The Generation Effect — Structural Learning&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:anthropic-comp&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Anthropic Salaries&lt;/em&gt; and related compensation write-ups. The reported $199K–$1.27M total-compensation band and the $1.12M–$1.38M H-1B base-salary filings. &lt;a href=&quot;https://www.levels.fyi/companies/anthropic/salaries&quot;&gt;levels.fyi&lt;/a&gt;, &lt;a href=&quot;https://www.levels.fyi/companies/anthropic/salaries/software-engineer&quot;&gt;levels.fyi (SWE)&lt;/a&gt;, &lt;a href=&quot;https://www.metaintro.com/blog/anthropic-1-3-million-ai-roles-salary-2026&quot;&gt;Metaintro&lt;/a&gt;, &lt;a href=&quot;https://ctaio.dev/en/salary/anthropic-salary/&quot;&gt;CTAIO&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-comp&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:openai-comp&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;OpenAI Salaries&lt;/em&gt; and related compensation write-ups. The Research Scientist and Member of Technical Staff bands and the reported retention-grant range. &lt;a href=&quot;https://www.levels.fyi/companies/openai/salaries&quot;&gt;levels.fyi&lt;/a&gt;, &lt;a href=&quot;https://jobsbyculture.com/blog/openai-compensation-2026&quot;&gt;JobsByCulture&lt;/a&gt;, &lt;a href=&quot;https://ctaio.dev/en/salary/openai-salary/&quot;&gt;CTAIO&lt;/a&gt;, &lt;a href=&quot;https://www.entrepreneur.com/business-news/how-much-openai-employees-make-salaries-685000&quot;&gt;Entrepreneur&lt;/a&gt; &lt;a href=&quot;#fnref:openai-comp&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:anthropic-careers&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Careers&lt;/em&gt;, Anthropic. Source of the two block quotes on hiring for demonstrated ability over credentials and on engineers doing research. &lt;a href=&quot;https://www.anthropic.com/careers&quot;&gt;anthropic.com/careers&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-careers&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-careers:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;sup&gt;2&lt;/sup&gt;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:anthropic-re-alignment&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Research Engineer / Scientist, Alignment&lt;/em&gt;, Anthropic (Greenhouse). Source of the “candidates need not have” quote and the “good fit” list. &lt;a href=&quot;https://job-boards.greenhouse.io/anthropic/jobs/4631822008&quot;&gt;job-boards.greenhouse.io&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-re-alignment&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-re-alignment:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;sup&gt;2&lt;/sup&gt;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:openai-re&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Research Engineer&lt;/em&gt;, OpenAI. Source of the quoted description of the role. &lt;a href=&quot;https://openai.com/careers/research-engineer&quot;&gt;openai.com/careers&lt;/a&gt; &lt;a href=&quot;#fnref:openai-re&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:interview-gate&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Sundeep Teki, &lt;em&gt;Anthropic Research Engineer Interview 2026&lt;/em&gt;. Account of the data-structures/live-coding interview process referenced in the discussion of interview gates. &lt;a href=&quot;https://www.sundeepteki.org/advice/anthropic-research-engineer-interview-2026&quot;&gt;sundeepteki.org&lt;/a&gt; &lt;a href=&quot;#fnref:interview-gate&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:anthropic-rsi&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;When AI builds itself&lt;/em&gt;, Anthropic. The claim that a model’s proposed next experimental step reaches parity with a human researcher’s in a substantial fraction of sessions. &lt;a href=&quot;https://www.anthropic.com/institute/recursive-self-improvement&quot;&gt;anthropic.com/institute&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-rsi&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:skill-atrophy&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;The Skill Atrophy Trap&lt;/em&gt;, TianPan.co, and &lt;em&gt;A Review of the Negative Effects of Digital Technology on Cognition&lt;/em&gt; (arXiv). The claim that unassisted judgment stalls at a ceiling set by the tool rather than the operator. &lt;a href=&quot;https://tianpan.co/blog/2026-04-19-skill-atrophy-ai-augmented-engineering&quot;&gt;tianpan.co&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/pdf/2603.10025&quot;&gt;arXiv&lt;/a&gt; &lt;a href=&quot;#fnref:skill-atrophy&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:unsourced-fluency&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;No source given.&lt;/strong&gt; Fluency as a cue for judged understanding, and the illusion of explanatory depth, are reported here from the general literature; no paper in this piece’s source list establishes either. Treat both as inference until they carry a citation. The argument does not rest on them alone — the METR result is the checkable version of the same claim. &lt;a href=&quot;#fnref:unsourced-fluency&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:metr&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;METR, &lt;em&gt;Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&lt;/em&gt;, and &lt;em&gt;AI Coding Tools Made Developers 19% Slower&lt;/em&gt;. The 19% slowdown, the 24% forecast and the 20% post-hoc estimate. &lt;a href=&quot;https://metr.org/blog/2026-02-24-uplift-update/&quot;&gt;metr.org&lt;/a&gt;, &lt;a href=&quot;https://letsdatascience.com/blog/developers-thought-ai-made-them-faster-the-data-said-otherwise&quot;&gt;letsdatascience.com&lt;/a&gt; &lt;a href=&quot;#fnref:metr&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:perception-gap&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;The Vibe-Check Protocol: Quantifying Cognitive Offloading in AI Programming&lt;/em&gt; and &lt;em&gt;Using Biometrics to Understand AI-Assisted Coding Performance and its Perception&lt;/em&gt; (both arXiv). Evidence that perceived speed and measured speed diverge and do not converge with experience. &lt;a href=&quot;https://arxiv.org/pdf/2601.02410&quot;&gt;arXiv (Vibe-Check)&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/pdf/2606.20598&quot;&gt;arXiv (Biometrics)&lt;/a&gt; &lt;a href=&quot;#fnref:perception-gap&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:metacog-laziness&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;Approximate support.&lt;/strong&gt; The term is used here for a mechanism the cited paper describes rather than names. &lt;em&gt;Mitigating “Epistemic Debt” in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts&lt;/em&gt; (arXiv). Source for the term and mechanism of metacognitive laziness under fluent AI output. &lt;a href=&quot;https://arxiv.org/html/2602.20206v2&quot;&gt;arXiv&lt;/a&gt; &lt;a href=&quot;#fnref:metacog-laziness&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:cognitive-debt&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;Approximate support.&lt;/strong&gt; A conceptual match rather than the paper’s own terminology. &lt;em&gt;Mitigating “Epistemic Debt” in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts&lt;/em&gt; (arXiv) and &lt;em&gt;AI-overdependence and human cognitive decline&lt;/em&gt; (ScienceDirect). Support for the cognitive-debt framing of comprehension deferred past the point of production. &lt;a href=&quot;https://arxiv.org/html/2602.20206v2&quot;&gt;arXiv&lt;/a&gt;, &lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S2451958826001764&quot;&gt;ScienceDirect&lt;/a&gt; &lt;a href=&quot;#fnref:cognitive-debt&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:unsourced-rpe&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;No source given.&lt;/strong&gt; The reward-prediction-error account of boredom, and the learning-progress framing of curiosity-driven agents, are asserted from background literature that is not in this piece’s source list. Treat this paragraph as inference. It motivates the design of the loops below; it is not evidence for them. &lt;a href=&quot;#fnref:unsourced-rpe&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:bjork&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Metcalfe &amp;amp; Bjork, &lt;em&gt;Desirable Difficulties and Studying in the Region of Proximal Learning&lt;/em&gt;. The retention-vs-performance tradeoff across generation, testing and varied practice. &lt;a href=&quot;https://www.columbia.edu/cu/psychology/metcalfe/PDFs/Metcalfe-BjorkVolSubmitFeb14Final.pdf&quot;&gt;columbia.edu&lt;/a&gt; &lt;a href=&quot;#fnref:bjork&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:undesirable-difficulty&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Undesirable Difficulty Effects in the Learning of High-Element-Interactivity Materials&lt;/em&gt; (PMC). The reversal of desirable difficulty when working memory is already saturated. &lt;a href=&quot;https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6099118/&quot;&gt;ncbi.nlm.nih.gov&lt;/a&gt; &lt;a href=&quot;#fnref:undesirable-difficulty&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:worked-example&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Desirable difficulties, worked examples, and completion problems&lt;/em&gt;, Frontiers in Psychology. The worked-example effect for novices and its reversal toward incomplete/completion examples with expertise. &lt;a href=&quot;https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.01623/xml&quot;&gt;frontiersin.org&lt;/a&gt; &lt;a href=&quot;#fnref:worked-example&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt; &lt;a href=&quot;#fnref:worked-example:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;sup&gt;2&lt;/sup&gt;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:taste&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Wenbo Pan, &lt;em&gt;Taste Is the Next Capability AI Will Crack&lt;/em&gt;. Discussion of research taste as the remaining, measured bottleneck. &lt;a href=&quot;https://www.wenbo.io/blog/taste-scaling/&quot;&gt;wenbo.io&lt;/a&gt; &lt;a href=&quot;#fnref:taste&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:generation-effect&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;The Generation Effect&lt;/em&gt;, Structural Learning. Support for generation over review as the stronger desirable difficulty. &lt;a href=&quot;https://www.structural-learning.com/post/generation-effect-active-learning&quot;&gt;structural-learning.com&lt;/a&gt; &lt;a href=&quot;#fnref:generation-effect&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:anthropic-fellows&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;em&gt;Anthropic Fellows Program&lt;/em&gt;, Anthropic (Greenhouse). Source for the program’s tracks, its selection on demonstrated capability, and the no-internships/12-month reapplication terms. &lt;a href=&quot;https://job-boards.greenhouse.io/anthropic/jobs/5023394008&quot;&gt;job-boards.greenhouse.io&lt;/a&gt; &lt;a href=&quot;#fnref:anthropic-fellows&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;
</content>
  </entry>
  
</feed>
