general 2026-09-27

Shopify CEO's 'slop grenades' warning: how to price the review cost of AI-generated work

Shopify CEO Tobias Lütke's warning about AI 'slop grenades' is a practical lens for SME teams weighing AI tools: the true cost of AI-generated work appears in the review hours it adds, not the volume it produces. This guide explains what Lütke said on The Knowledge Project podcast, why researchers call the pattern 'workslop,' and what engineering telemetry reveals about review burden. It then sets out quality gates — evidence checks, human-in-the-loop review, and stake-matched verification — plus a simple method for small teams and AI tool buyers to measure whether a writing, coding, or automation tool saves or adds review time before committing.

What the "slop grenades" warning actually means

On "The Knowledge Project" podcast, released on Tuesday, Shopify CEO Tobias Lütke said AI can make work more difficult by enabling what he called "slop grenades" — AI-generated material passed along without proper evaluation. "We call those 'slop grenades' that people toss at each other," he said, and called it "definitely a bad thing". His concrete example was an unnecessarily long AI-written email that forces the recipient to run it through another large language model just to summarise it.

The issue is the missing review step, not the existence of the draft. Lütke said the tools make it easy to pass along work that has not been properly evaluated, and that AI is least valuable when it adds slop to a coworker's desk. He had earlier told Shopify staff that using AI was "a baseline expectation" and that they should first prove they "cannot get what they want done using AI" before asking for more resources. The tension is that the same push for AI use now produces unchecked output that creates more work for the people receiving it.

Lütke was explicit that this is not a verdict on AI as a whole. He said AI is most valuable when it makes people's thinking clearer and more concise, and that "humans take responsibility" — with machines helping people take more responsibility because they inform them better. That distinction is the spine of this article: the problem is unreviewed output, not the tool.

"Workslop" is a named pattern, not a complaint

Researchers coined the term "workslop" for exactly the phenomenon Lütke described: polished-looking AI output that drags down productivity because it needs revision. What separates workslop from ordinary low-quality human work is that it looks legitimate on the surface while missing the components that make it useful. Survey respondents cited well-structured emails containing broken links and code that was more complicated than necessary as typical examples.

A survey of 1,150 office workers found that 40 percent had encountered workslop — AI-generated content that appears professional but lacks actual substance. Duolingo CEO Luis Von Ahn described a parallel case, saying that when scaling content with AI, roughly 20% of items were "pure slop" and had to be caught before reaching users. The shared lesson across both executives is that output volume from AI does not equal delivered value; the review layer is where value is either protected or lost.

The review burden shows up in the numbers

The cost of unreviewed AI work is more than anecdotal. In a 2026 developer survey of more than 1,100 respondents, 72% of developers who had tried AI coding tools said they now use them every day, yet 96% said they do not fully trust that AI-generated code is functionally correct. Only 48% said they always check AI-assisted code before committing it. Developers still spend about 24% of their week on toil work — tasks that sap productivity — with the burden shifting toward correcting and rewriting unreliable AI-generated code.

Telemetry across thousands of engineering teams found that under high AI adoption, the median time to first pull-request review rose 156.6%, average time in review rose 199.6%, and median time in review rose 441.5%. Agentic code reviews surged in 2026, with 25% of pull requests reviewed by an AI agent, up from 0% in 2025, while review times increased nearly 200%. In the same high-adoption settings, 31% more pull requests were being merged without any review, human or agentic. The replacement cost of a senior software engineer was placed at $150,000 to $300,000 in 2026, factoring in recruiting, ramp time, and lost institutional knowledge.

These figures come from research published earlier in 2026 and should be read as background on the verification bottleneck, not as this week's news. They help explain why Lütke's warning lands: the constraint has moved from writing the work to verifying it.

Shopify's own AI agent shows the other side

Lütke's warning sits next to Shopify's aggressive internal AI adoption, which shows AI can carry real load when review is designed in. He said Shopify's internal agent, River, handles a large share of the company's production code pull requests — possibly as much as half. River is an AI agent that lives in Shopify's company Slack, where employees mention it in a public channel to read code, run tests, open pull requests, query the data warehouse, and inspect production traces.

The design choice that matters for review cost is that River only works in the open, with no direct messages, so every conversation becomes a public Slack transcript other employees can read. In one reported 30-day window, 5,938 Shopify employees worked with River across 4,450 Slack channels, and in one week River opened 1,870 pull requests in the main monorepo, with about one in eight merged pull requests coauthored by it and reviewed by humans. River's merge rate improved from 36% to 77% over two months, attributed not to a model swap but to people watching River work, noticing where it got stuck, and feeding corrections back.

The lesson for SME teams is not to build a River of their own. It is that AI output scales safely when the review step is visible, shared, and repeated — the opposite of a private, unreviewed "slop grenade" landing on a colleague's desk.

Quality gates that keep AI useful

The practical answer to review cost is to make verification a defined step rather than an afterthought. One framework organises trust in AI output into five pillars: evidence, grounding, human review, correction, and governance. The core rule is that an AI output can be trusted when you can see the evidence behind it, verify it against your own context, and a human still owns the final decision.

A repeatable workflow is to match review depth to the stakes: a rough brainstorm gets a light look, while a spec an engineer will build from needs a hard review. Human-in-the-loop review means a person reviews and approves AI output before it drives a real decision, with AI doing the first pass and a person validating and deciding. Correction workflows treat feedback as quality control — a correction fed back makes the next output better, much like a good bug report improves software. One report notes that nearly two-thirds of respondents say their organizations have not yet begun scaling AI across the enterprise, which suggests the teams that do scale it are the ones that have defined where and when model outputs need human validation.

For teams piloting any AI tool, a quality-control guide suggests starting small with pilot programs to refine processes before scaling, setting clear KPIs, and keeping detailed records that show where human intervention was required. The principle transfers directly: prove the review load is manageable on a narrow task before widening the tool's reach.

How SME teams can measure review cost before buying

For an Australian or cross-border SME evaluating an AI writing, coding, or automation tool, the question to answer is not "how fast does it produce?" but "how many review hours does it add versus save?" A useful test is to run the tool on one real task and time both the generation and the verification separately. If the recipient spends longer checking the output than they would have spent producing it themselves, the tool is adding slop rather than removing work.

They can also check whether a candidate tool offers the guardrails that reduce review burden — summary-first outputs, diff limits, source or citation links on factual claims, and a clear marker of what was machine-generated.

The decision framework is straightforward. Test the tool on a representative task, measure generation time and review time, confirm a human keeps the final call, and only then widen access. AI is not inherently productivity-negative; unreviewed AI output is.

Reader questions on AI review cost

What exactly did Shopify's CEO say about "slop grenades"?

Tobias Lütke said on "The Knowledge Project" podcast that AI can make work harder by enabling "slop grenades" — AI-generated material passed along without proper evaluation — and that AI is least valuable when it adds slop to a coworker's desk. He said "humans take responsibility".

What is "workslop" and how common is it?

"Workslop" is the term researchers use for polished AI output that looks legitimate but lacks useful content, dragging down productivity because it needs revision. A survey of 1,150 office workers found 40 percent had encountered it.

Does AI-generated code actually increase review time?

What quality checks stop AI output becoming slop?

A five-pillar framework uses evidence, grounding, human review, correction, and governance, with review depth matched to stakes and a human owning the final decision.

How should a small team test an AI tool before buying?

Run it on one real task, time generation and review separately, and only widen access if review time stays below the time saved, with a human keeping the final call.

Sources

Faros.ai, "How AI-Generated Code Is Increasing Code Review Burden" (2026-05-21). Sonar Community, "2026 State of Code Developer Survey" (2026-01-08).