Just because it doesn't materially impact task success rates does not mean it's not useful. I use that file to give my agent information about the environment (operating system, architecture, command-line utilities I have installed), as well as how to do certain things (such as using `uv` for running Python when needed). I want my agent to work how I do, so I also tell it things like my preferred version control system (jj), my preferred implementation languages for things like shell scripts (zsh), and other things like that. I care more about how the work is done than the final result. It still slips up sometimes, but on the whole I think it works alright. It's not like it fixes tasks that wouldn't be completed at all, but it does help my satisfaction with how they were completed, as well as with the final result. Otherwise, I'd have to do a lot more manual cleanup.
I just have it comment on scripts, I've now just started for an app super-run.sh consisting of this pipeline: dev, test, build, e2e, deploy.
All the scripts are designed to bail quickly. They all create a log in the background instead of blocking. They're all annotated with the necessary comments to keep it on path and locate files, track pratfalls, etc.
So any entrypoint to whatever I'm doing typically starts with one of these scripts. It's heavy handed but orientating your agent for the specific task is better than just dumping a whole set of context that will be ignored if it has nothing to do with your very next command.
Telling it how to do a pull request isn't going to help if you're trying to debug a technical issue.
I haven't really been doing this yet, but for what it's worth, I think explaining how to use tools doesn't really belong in your agents.md. I think it's better to use it to explain the what tools are available (information about the environment, like you said), and then provide details on the how in skills files. That way you're not bloating your context with how to use `uv` in a session where you're not doing anything Python related. At worst, you're just saying that `uv` exists on your system.
It mentions that a chatbot generated AGENTS.md does nothing, which makes sense.
I added one when it kept making the same mistake and using things from the wrong library version making compiler errors, adding in the common mistakes pre emptively.
On my last project, it kept trying to use the system python instead of the project's virtual environment. It also kept using the wrong build tool. Both things wasted considerable tokens because the agent got sidetracked trying to understand why it could not run the tests - and that repeated on each new session.
This is what the majority of mine look like as well. Just simple instructions for things I need to do repeatedly.
I have not even really needed a formal memory system. If I see an error happen more than once, I just say "hey add a note on this to agents.md". Tends to be verbose but overall works quite well for the projects I am doing.
Yeah, I feel like what these files do is not that difficult to understand; it's just context that the agent will pick up and use pretty much the same as any other context it has. No, it won't deterministically prevent things with a static check, but it will work about as well as just manually telling the agent "don't do X" as part of the prompt. The fact that it's a "mostly works" mechanism rather than a "guaranteed to always work" mechanism is pretty much the same experience that using an LLM gives in general, and while that requires a bit of thought about how to use it, it's still good enough to be useful for a lot of things.
Before LLMs, I found "mostly works" systems like this to be incredibly sketchy and not worth using. The main thing I've had to learn in this past year is that what I thought was an ironclad rule turned out to be only a heuristic that was useful before but not always helpful, because empircally as much as I might find the lack of determinism jarring, in practice these tools are genuinely good enough at what they do to be worthwhile to use, as long as you're making sure not to use them in ways that the occasional failure costs more than just some wasted time.
This research is not current, and is based on a flawed premise - generated agents.md are used. There isn’t a hand curated agents.md that has some knowledge passed to it that is withheld then measured against.
I think there is overall something here for current Claude which is there appears to be a hierarchy of conformance that breaks progressive disclosure and the utility of skills. It seems to honor the system prompt, user instructions, tool call results, and dead last skills. It applies a large amount of discretion as to whether to honor what skills say in the imperative and progressive disclosure seems to have at best a 20-30% recall. Other models like codex gpt 5.6 seem to be the exact opposite and slavishly adhere to the Agent/skills/plugins, to the point of being wasteful and dangerous. It feels clear there’s a tension being RL’ed around between conformance and skeptical behavior that neither has quite found the balance for yet, and is almost certainly an over constrained problem. I just find it funny Anthropic is the one you can’t trust with your wallet while OpenAI does precisely what you and your harness tell it to.
But this “science” and its editorializing are based on flawed techniques, don’t lead to the conclusion let alone the editorialized extrapolation, and are l
Never is a strong take (for an easy to find example, the Bun port posts had a ton of comments about the interests of the author rather than the results), but even if we switch it to a "much less commonly" interpretation of the phrase I'd say most pro AI articles talk about what someone/some group has been doing with it and that's what gets the majority of the skepticism goes. On the other side, most anti-AI articles are about what the author sees, so they they tend to be the focus of the skepticism instead.
Still, I agree it doesn't make much sense to only talk about who's writing the content instead of the content. Perhaps, generously, the other comments covered their opinions on fhat part already.
I have in mine to not do any write operations with git, to only use the CLI to explore the history and the current state. I don't trust my agent to commit and push on my behalf. Using my agents.md has worked well for that kind of thing.
I have no historical context (heh) for this person, but as someone who uses LLMs regularly, I feel like there are reasonable takeaways from the article. Namely:
- Don't vibe you agents.md file, it won't capture any intuition about the project that the model doesn't already have.
- Keep your agents.md file short. Long ones mostly bloat context for minimal difference in behavior.
- Writing for a human audience is probably better. Any LLM can read docs made for humans anyways.
All the scripts are designed to bail quickly. They all create a log in the background instead of blocking. They're all annotated with the necessary comments to keep it on path and locate files, track pratfalls, etc.
So any entrypoint to whatever I'm doing typically starts with one of these scripts. It's heavy handed but orientating your agent for the specific task is better than just dumping a whole set of context that will be ignored if it has nothing to do with your very next command.
Telling it how to do a pull request isn't going to help if you're trying to debug a technical issue.
I added one when it kept making the same mistake and using things from the wrong library version making compiler errors, adding in the common mistakes pre emptively.
A simple instruction in AGENTS.md fixed that.
I have not even really needed a formal memory system. If I see an error happen more than once, I just say "hey add a note on this to agents.md". Tends to be verbose but overall works quite well for the projects I am doing.
Before LLMs, I found "mostly works" systems like this to be incredibly sketchy and not worth using. The main thing I've had to learn in this past year is that what I thought was an ironclad rule turned out to be only a heuristic that was useful before but not always helpful, because empircally as much as I might find the lack of determinism jarring, in practice these tools are genuinely good enough at what they do to be worthwhile to use, as long as you're making sure not to use them in ways that the occasional failure costs more than just some wasted time.
Either way, you get no visibility or predictability.
Good luck.
That's fine. Do what you do. But don't read this article as any sort of science. It's massively opinionated rage-bait.
[0]: https://www.patreon.com/davidgerard
I think there is overall something here for current Claude which is there appears to be a hierarchy of conformance that breaks progressive disclosure and the utility of skills. It seems to honor the system prompt, user instructions, tool call results, and dead last skills. It applies a large amount of discretion as to whether to honor what skills say in the imperative and progressive disclosure seems to have at best a 20-30% recall. Other models like codex gpt 5.6 seem to be the exact opposite and slavishly adhere to the Agent/skills/plugins, to the point of being wasteful and dangerous. It feels clear there’s a tension being RL’ed around between conformance and skeptical behavior that neither has quite found the balance for yet, and is almost certainly an over constrained problem. I just find it funny Anthropic is the one you can’t trust with your wallet while OpenAI does precisely what you and your harness tell it to.
But this “science” and its editorializing are based on flawed techniques, don’t lead to the conclusion let alone the editorialized extrapolation, and are l
Still, I agree it doesn't make much sense to only talk about who's writing the content instead of the content. Perhaps, generously, the other comments covered their opinions on fhat part already.
Knowing he did an AI hate interview with Tante just solidifies this: https://pivot-to-ai.com/2026/08/21/tante-on-ai-when-this-thi...
Just looking through his Mastodon reposts (neovim is "fascist software" if you didn't know already!) just leaves me shaking my head.
Incredible how much of an audience you can get by just being anti "the latest hype".
- Don't vibe you agents.md file, it won't capture any intuition about the project that the model doesn't already have.
- Keep your agents.md file short. Long ones mostly bloat context for minimal difference in behavior.
- Writing for a human audience is probably better. Any LLM can read docs made for humans anyways.