II run a pipeline daily that searches the web, curates what it finds, and publishes a page. Six of its eleven steps call a model. Five never do — and those five are the ones that make it safe to leave running.
Model: searching each topic, extracting structured items, ranking and picking the lead, reviewing the result, writing a line of commentary.
Plain Python: date and history, the rules gate, rendering, uploading, verifying the live URL afterwards.
The gate is the argument. Blocked domains, duplicate URLs, nothing republished within 7 days, a hard item cap. All four started as lines in a prompt, and all four got promoted to code — because "the model follows this most of the time" is fine while you're watching and useless on a schedule. Over a year of unattended runs, "most of the time" is a stack of small embarrassments nobody was there to catch.
The split I've landed on: judgement goes to the model, invariants go in code. Which of two stories is bigger is judgement. Whether this URL ran last Tuesday is a set lookup, and it should never be anything else.
That has a price and I'll name it. My image selection is pure code — width, aspect ratio, filename blocklist — and it quietly rejected real editorial art for weeks, because CMSs serve thumbnails and a 480×320 derivative of a good illustration fails a width check. The rule was correct and the outcome was wrong. That's the trade: code gives you rules that always run, and rules that are confidently wrong in ways nobody notices.
I still think it's the right trade. Blunt and predictable beats sharp and occasionally absent.
So where's your line? Specifically: what did you move out of code because deterministic turned out too blunt? That direction gets argued a lot less than the other one, and I suspect it's where the interesting answers are.
LangGraph pipeline, running daily.
Code: https://github.com/ravi-labs/agentic-newsroom
Write-up: https://medium.com/@rkanagasikamani/the-newsroom-that-writes-itself-8c0160f68aac
Source: r/LangChain · by /u/Mysterious_Hunter_92