METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Karpathy's LLM coding critique, distilled into a single CLAUDE.md file

A guideline file that curbs Claude Code's habit of making assumptions and bloating code has hit GitHub's trending page

Karpathy's LLM coding critique, distilled into a single CLAUDE.md file

Image: METAL

Summary

  • The multica-ai repository built a CLAUDE.md guideline based on Andrej Karpathy's critique of LLM coding, and it recently landed on GitHub's trending page.
  • The core CLAUDE.md file is built around four principles: no assumptions, minimum code, minimum edits, and success criteria.
  • Installing it as a Claude Code plugin or a Cursor rules file lets you apply the same guidelines across multiple projects.

Karpathy's observation, boiled down to one file

Andrej Karpathy pointed out a recurring habit of LLM coding agents in a post on X. Building on that observation, the multica-ai repository put together a set of guidelines you can plug into Claude Code, and posted it under the name "andrej-karpathy-skills." The repository, created in January 2026, recently climbed GitHub's trending page. The core guideline lives in a single CLAUDE.md file, but the repository also includes a README, CURSOR.md, SKILL.md, a plugin, and Cursor rules files to help with installation and integration with other tools. CLAUDE.md contains four principles designed to address the three problems Karpathy identified.

Three habits Karpathy called out

Karpathy argued that LLMs repeat three habits when writing code. "Models make the wrong assumption on the user's behalf, then just proceed without checking." In other words, they don't manage confusion, don't ask for clarification, and don't surface contradictions or trade-offs. The second is over-engineering: they stretch a job that should take 100 lines into more than 1,000, and never clean up dead code. The third is side effects — changing or deleting comments and code unrelated to the request, without fully understanding them.

GitHub 저장소 andrej-karpathy-skills 메인 화면과 파일 목록, README 일부가 보임
이미지: GitHub · multica-ai

Four principles in CLAUDE.md

The repository lays out four principles addressing these three problems.

PrincipleKey phraseDescription
No assumptions"Don't assume. Don't hide confusion. Surface tradeoffs."Instead of quietly settling on an interpretation and moving ahead, the model surfaces uncertainty and trade-offs first
Minimum code"Minimum code that solves the problem."A principle against over-engineering, with the standard: "if a senior engineer would call it excessive, simplify it"
Minimum edits"Touch only what you must. Clean up only your own mess."Keeps the model from touching code unrelated to the request, and requires every changed line to be traceable back to the request
Success criteria"Define success criteria. Loop until verified."Turns imperative instructions into verifiable goals, letting the model iterate on its own until it confirms the criteria are met

The fourth principle connects to another of Karpathy's observations: give a model a clear goal and it will iterate on its own to meet it, but give it a loose standard like "make it work" and you'll end up going back and forth with it repeatedly.

GitHub 저장소 CLAUDE.md 파일 내용 일부가 코드 가이드라인 형태로 표시됨
이미지: GitHub · multica-ai

How to try it

The multica-ai/andrej-karpathy-skills repository walks through two installation methods. The recommended approach is the Claude Code plugin. Adding the marketplace inside Claude Code installs the guidelines as a plugin, so the skill applies to every project you open in Claude Code, not just one specific project. The second option is Cursor. The repository includes a .cursor/rules/karpathy-guidelines.mdc rules file, so the same principles apply as soon as you open a project in Cursor. Setup instructions for Cursor are laid out separately in the repository's CURSOR.md document.

If your project already has a CLAUDE.md file, you can simply merge these four principles into it, and if there are project-specific rules you need to follow, you can add a section below the guidelines. The repository notes that the guidelines aren't meant to weigh down trivial tasks like fixing a typo — the focus is on reducing mistakes that are hard to undo.

Creating a rules file and sharing it across multiple coding agents has the advantage of letting a team reuse its working principles instead of tying them to one specific model.

The post announcing the repository also introduced Multica, a separate open-source project by the same creator. It's described as a platform for running and managing coding agents using reusable skills, and it's a separate project from the CLAUDE.md repository.

Editor's take

What makes this repository interesting isn't that it turned Karpathy's personal observations into rules verbatim. It's that it shows just how identical the problems are for anyone using coding agents in real work. Whether it's Claude Code or Cursor, switching models doesn't change the complaints — "it quietly assumes and pushes ahead" or "it bloats 100 lines into 1,000" come up everywhere. This isn't a bug in one particular model — it's closer to a habit shared across LLM coding agents as a whole. What this case demonstrates is that a single guideline file can correct that habit to a meaningful degree.

Put a coding agent of similar scale to work in a real project, and the conclusion tends to be the same. If you hand over a loose goal like "just make it work," the agent pushes ahead with its own interpretation instead of asking questions — and reviewing the result ends up taking even more time. Give it a verifiable standard instead — "success means this function produces this output for this input" — and you can watch the agent iterate on its own to meet it. The fourth principle in this CLAUDE.md hits exactly that point.

For teams writing code with Claude Code or Cursor, adding a guideline file like this to a project costs almost nothing. But to actually see results, each team needs to first identify the mistakes that keep recurring for them — not cleaning up dead code, touching unrelated files — and then tailor the principles to address those specific issues. Simply copying and pasting someone else's rules only gets you halfway there.

In the coming weeks, we're likely to see more sharing of CLAUDE.md files and skill files that distill one person's observations or a team's experience this way. If the open-source community settles into a pattern of trading rules files instead of prompts, it could well lead developers like Anthropic to eventually fold proven rules into their default system prompts.


Correction (2026-08-24) — The original article described these guidelines as newly released. The repository was actually created on January 27, 2026. The lead and summary have been corrected.

Comments