The pattern that only shows up in the batch
51 lifestyle images, generated one category at a time. Every one of them passed on its own; the page they made together didn't.
The test
I gave ChatGPT a spreadsheet: 51 category names, each with a link to that category's product landing page on a live ecommerce site. The brief was deliberately thin: visit each page, look at the products listed there, and produce a lifestyle image for the category showing humans using those products.
No further instruction. No style guide. No rule about what "using" the product should mean in different categories. That was the point: I wanted to know what a model does with a real, unglamorous merchandising task when it's given nothing beyond the task itself.
For most of the 51, it did the job. Real products, pulled from the actual page, composed into a plausible lifestyle shot. That's not nothing; it's a genuinely capable piece of automation, and worth saying so plainly before getting into where it broke.
Where it broke, and why each break is a different kind of gap
Wrong scale. Some products came back sized incorrectly relative to the humans or the setting around them. Each image was solved on its own: nothing carried a shared sense of proportion from one generation to the next.
Wrong orientation. Similarly, some products faced the wrong way, or sat in the frame in a way that wouldn't survive a human reviewer's first pass. Same root cause: no consistent rule enforced across the batch, because nothing told it there should be one.
Missing the evidence of use. This is the sharpest one. The Oktoberfest category came back with drinking glasses in shot, with no drink in them. Plates, in other categories, with no food on them. The model rendered the object correctly. It didn't reason about what "in use" implies once you're looking at the whole scene: a glass mid-use has something in it. That's not a missing fact so much as a missing narrative rule: the kind of thing a human stylist would apply without being asked, because it's obvious once you're looking at the finished scene rather than generating it prop by prop.
Human-density drift across the page. Looked at individually, none of the 51 images seemed wrong. Looked at together, as the actual page a customer would scroll through, most images had settled on roughly four people in shot. Fine once, heavy at scale: the page reads as visually dense and repetitive in a way no single image reveals.
Why the fourth one matters most
The first three are all versions of the same thing: the model had no context about its own output. The fourth is different in kind. Each of those 51 images was generated correctly relative to its own brief. The problem only exists at the level of the page, the set of 51 taken together, which nothing in a single-image generation loop is positioned to see.
That's worth naming as its own thing, distinct from the failure patterns already documented on this site. Everything in How Context Debt shows up is about a system missing a fact it needed to do its job. This is about a batch missing context about the other items in the batch: call it Batch blindness, the difference between item-level context and page-level context. A model can be given everything it needs to get one image right and still produce a page that reads as wrong, because "right" at the page level is a judgement about the set, not about any single member of it.
The fix isn't "add more context" in general. It's specific, the same way every fix on this site is specific: define the distribution before the batch runs. If a page needs variety in how many people appear per image, that's a rule to set upfront, a percentage split, agreed once, applied consistently, not something to catch on review and patch image by image.
The fix, in practice: Oktoberfest
Same scene both sides: three people toasting outdoors at a table dressed for Oktoberfest. In the uncorrected version, two of the three glasses are empty; only the centre glass holds beer. In the corrected version, all three glasses are filled. Nothing else in the frame changes: same people, same poses, same table, same decorations. The only difference is the one that was specified.
The original instruction was, in effect, "make a lifestyle image of people using these products." The corrected instruction added one explicit rule: if drink glasses are in shot, they should contain an appropriate drink; the same narrative logic applies to plates and food, and to any other product context has already established.
Nothing else changed. Same model, same category, same source products. The only difference is a single sentence stating a rule a human wouldn't have needed spelling out, which is exactly the pattern this site exists to document: the output didn't get better because the model improved. It got better because someone decided, in advance, what the model needed to know.
One thing the fix didn't fix
Small confession, while I'm being precise about what changed: the first corrected pass still wasn't right. The middle glass was fixed and filled, but the woman's hand wasn't actually gripping it, just hovering nearby in solidarity. That's a second issue caught in the same image, and it's a genuinely different kind of problem: not a missing rule about context, but the model's grip on hand-object anatomy breaking down under a specific pose. Worth flagging here rather than quietly fixing and moving on, because it's the first sighting of what looks like its own pattern; one this site hasn't named yet, and will when there's enough evidence to do it properly.
Where this sits
This ran through ChatGPT's browser interface, not Signal & Flow: worth being precise about that. It also wasn't a controlled demo built to prove a point; it surfaced while doing a real task with a deliberately vague brief. That's part of why it's useful: it's evidence of the pattern occurring on its own, not evidence manufactured to fit an argument already made.
Frontmatter: content/research/the-pattern-that-only-shows-up-in-the-batch.md
title: The pattern that only shows up in the batch slug: the-pattern-that-only-shows-up-in-the-batch order: 4 summary: 51 lifestyle images, generated one category at a time. Every one of them passed on its own; the page they made together didn't. stat: '51' stat_label: images in one batch, and the pattern that broke it only became visible once they sat side by side
If this way of thinking is relevant to a problem you're facing, I'd be glad to talk it through.
Start a conversation