Methodology
How the eval gets run
The name is a promise, so here's the machinery behind it. Every article on this desk goes through the same gauntlet before it ships — and when a piece can't survive the gauntlet, it doesn't ship. That's the whole system.
Two sources or it doesn't ship
Every load-bearing claim — a price, a benchmark, a launch date, a quote — gets verified against at least two independent sources. Can't confirm it? It gets cut or stated as unverified. We'd rather be general and right than specific and wrong.
Dead links are fabrications
Every citation in every article is machine-checked: the URL has to actually resolve. A citation that 404s is treated exactly like a made-up source, because functionally it is one. This runs automatically on every piece, every time.
The adversarial pass
Before publication, a second, separate review pass gets one job: try to refute the article. It hunts for names, numbers, dates, and quotes it can't independently confirm. Anything flagged pulls the piece out of the publish queue and into human review. The default is to flag — false alarms are cheap, published fabrications are not.
No invented receipts
First-person testing claims ("we ran this," "it took 40 minutes") appear only when somebody actually did the thing. If we haven't touched a tool, the article says what the evidence shows — it never cosplays hands-on experience. This rule exists in writing and is enforced at review.
Both columns, always
Every tool gets a W column and an L column. A review that only finds strengths is an ad; a review that only finds flaws is a grudge. If we can't say where a tool wins and where it loses, we haven't finished the eval.
Exact numbers or honest silence
When we have a real number — a price, a context window, a rate limit — we print it exactly. When we don't, we don't round one into existence. "About $20/month" is honest. "$19.83 in our testing" is only honest if there was testing.
The pipeline, honestly
This desk publishes daily, and no, one human isn't hand-typing every word. Our research pipeline drafts with AI using live web search for grounding — then every draft is forced through the checks above: the two-source rule, machine-verified citations, and the adversarial pass that tries to tear the piece apart before you ever see it. Drafts that fail get held for human review instead of published.
The editorial standard belongs to Micah Berkley — The AI Mogul — who built the pipeline, wrote its rules, and signs the verdicts. AI does the legwork. The judgment, and the accountability, stay human.
Think of it this way: we run evals on AI tools, so we run our own newsroom the same way — test the output, verify the claims, publish the receipts. If that standard slips, call us on it.