Files and folders vs. agents
The infrastructure debate over filesystems and vector databases is the same argument as context design, one layer down the stack.
The claim
There's a debate among AI engineers that looks like it's about infrastructure: files and folders or agents and frameworks, filesystems or vector databases, readable markdown or unreadable embeddings and similarity search. People treat it as a question of tooling. It's actually a disagreement about context. The argument has just moved down from the prompt level into the architecture underneath.
The pitch for complexity
For a couple of years, the assumed shape of a serious AI system was: chop documents into chunks, embed them, drop them in a vector database, and let the model pull back whatever's semantically closest to the question. That's a reasonable answer to a specific problem, a static pile of documents and a model with a limited context window. Retrieval-augmented generation earned its place solving that. But an agent doing real work doesn't ask one question and leave. It reads a file, notices a reference, follows it, reads the next file, runs something, checks the result, and adjusts. That's closer to how a person navigates a project, and it turns out plain files and sensible folder names hold up better under that kind of exploring than a database of disconnected chunks does.
The folder is the agent
Anthropic's engineering team makes the comparison directly: people don't memorise everything, they lean on external organisation like file systems, inboxes and bookmarks to find things when they need them, and agents can work the same way. Their example is a good one. A file called test_utils.py means something different in a tests folder than it does in src/core_logic. Nobody wrote that meaning down; the location carries it. Take the idea one step further and the agent isn't really the model at all. It's a project folder: a file telling the AI what it's for, a set of skill definitions and a history of accumulated context, with a general model applied to it. Change the folder and the same model behaves like a different specialist. The specialism isn't in the model. It's in what the folder gives it.
Why this is the same argument
Which means the "files vs. agents" debate isn't really about files. It's about whether the work of curation happened before the model was asked to do anything, or whether everyone's hoping a cleverer retrieval mechanism can substitute for that work never having been done. It can't. A vector database full of unfiltered chunks is context debt with an API in front of it. It looks like infrastructure. It's actually an unmade decision about what matters, deferred to query time and hoping the maths sorts it out.
Files make you show your work
A folder of files forces decisions. You have to decide what belongs in it, what doesn't, how things relate to each other, and what names will make sense to whoever comes next. A messy folder is obvious. A vector database can hide the same mess behind similarity scores: it never makes you decide what's relevant in advance, and it always returns something that looks right. A plausible wrong answer is much harder to spot than an empty folder. A model working from incomplete information doesn't ask for clarification the way a new employee would; it produces a confident answer and moves on. A retrieval system that hands it the wrong chunk fails the same way, somewhere nobody is likely to look.
The honest caveat
The "just use files, vector databases are dead" take going around is overstated. At scale, with fuzzy queries and concurrent access, structured retrieval still earns its place, and most production systems use a mix of both. I'm not arguing for one kind of infrastructure. I'm pointing at the pattern underneath: whether you use files or vectors, the quality of an AI system depends on decisions made before anyone picks a storage method. What matters, how it's structured, what's left out. Files and folders are just the format that makes it hardest to pretend that work has been done. That's why they're the preferred option right now.
--- date: 2026-08-20 updated: 2026-10-03 title: Files and folders vs. agents slug: files-and-folders-vs-agents summary: The infrastructure debate over filesystems and vector databases is the same argument as context design, one layer down the stack. stat: 1 layer stat_label: how far down the stack the same argument about context repeats ---
If this sounds like your business, the audit is the quickest way to find out what your AI is missing: one workflow, one 90-minute session, and a written brief on what to fix first.
Book the auditPrefer to talk first? Get in touch — no obligation.