Hakkan
Built a content research tool to fight AI slop — content is created strictly from a researched report, real articles and real conversations, and the tool shows you when a number is derived rather than researched.
Hakkan (発刊, “to publish”) is a research-first content tool. Give it a topic and it reads the actual conversation — Reddit, TikTok, X, YouTube, the open web — then hands you a visual report with receipts, and helps you build content from that research in your own voice. This is the story of why it exists and what it took to make “no slop” true rather than a tagline.
The problem: content stopped knowing anything
After more than 12 years in marketing and content, I watched the medium slide from insightful to what everyone now calls AI slop — fluent text that reads fine and knows nothing, generated from a blank page and a guess. The tools caused it: every AI writer starts from nothing and asks the model to fill the void.
Hakkan inverts that. It starts from research — thousands of real articles, posts and conversations matched to your question — and treats that research as the source of truth. Content is built strictly from the report: the model is never the source, the research is. Whatever the model cannot do, you add your human touch to, on purpose.
The name carries the philosophy: Hakkan (発刊) is Japanese for “to publish”, and said aloud it echoes “harken” — to listen closely. Listen first, then publish. The design borrows from paper and the markers we abused during study, because the product’s whole argument is that the oldest publishing values still apply.
How it works: research → report → your voice
You give Hakkan a topic or a question. It fans out across social platforms and the open web, gathers the conversation, filters it for relevance, and builds a visual report: themes categorised, sentiment weighed, angles ranked, and every quote cited to the person who said it. From there you can export the research, or generate content from it — in a persona trained on your own writing, at a format and length you choose.
The balance is deliberate. Automation does the reaching and consolidating; the taste, the opinions and the final voice stay human. Some work is left manual on purpose, because the mistakes and the opinions are the part of content that connects.
What one report is actually made of
A real shipped report, on 2026 World Cup commercialisation: 965 items cited, 903 of them real voices — comments, posts and transcripts from people, not publications.
| Source | Items cited | What changed |
|---|---|---|
| Open web | 900 | Articles and pages, reached through search APIs. |
| Google News | 30 | The press layer, kept separate from opinion. |
| 21 | Where the arguments actually happen. | |
| YouTube | 5 | Transcripts, so video voices are quotable. |
| Perplexity | 4 | Synthesis, weighted low on purpose. |
| 4 | ||
| 1 |
The mix is chosen per question, not fixed: a travel question reaches TripAdvisor, a developer question reaches Hacker News, without a code change. Depth is a user choice, not a limit the tool invents.
The hard part: teaching the filter to value people
The promise is “what people actually said”, and the first version of the evidence filter quietly betrayed it. A verbose article restates the topic in its own headline, so it scored high; a real reply — “Why need a nanny if I won’t have a job” — is short, oblique and contextual, so it read as off-topic and died. The filter was killing exactly the material the product sells.
The fix was a rubric that judges a comment as a comment: replies answer the thing they reply to, not your search query, so relevance is scored in context, and only spam, bots and meta-chatter score zero. Rewritten on principle, then measured once — not tuned to the target.
Human voice in the evidence, measured at every stage
The bar was set at 40% of cited items being real human utterances. Getting there took four attempts — including catching our own test harness lying to us.
| Stage | Voice ratio | What changed |
|---|---|---|
| First measure | 12% | Later invalidated: a test flag was replaying cached data instead of searching live. Caught, documented, runs deleted. |
| Live streaming | 29% | Real runs, streamed comments arriving. Better — and still losing voices in the filter. |
| Facts added | 19% | More article evidence diluted the voices. The filter was the bottleneck, now provable. |
| Filter rewritten | 57% | Comments judged as comments: 81 real voices cited in the acceptance run, against the 40% bar. |
The 12% row stays in this chart deliberately. The measurement was taken on replayed fixture data a flag had switched on, and the moment that was discovered it was written down and the affected runs were deleted. A tool that sells honesty has to be built by a process that practises it.
Honest by construction
“No slop” is enforced in code, not tone of voice. The product refuses whole categories of fabrication that competing tools happily ship:
- 3-way
- Every number is classifiedGrounded in the research, drawn from your own writing, or derived by the model — and only the third is flagged for you to judge. Flagged, never blocked: only the author knows which figures they stand behind.
- 0
- Virality predictionsRefused outright. With no outcome data to train on, a prediction is a made-up number beside real citations. Hakkan shows what did break out, measured against each platform's median.
- “of the N voices here”
- Every claim is scoped60% of the voices in a report being frustrated is measured and true; “60% of people” is neither. The copy states the scope, never an apology.
Same rule inside the business: the tool was built against a hard cost-per-run ceiling enforced in code, so the research depth users get is sustainable for them and for us — priced to be worth it on both sides, without a single invented limit.
The screens
What I'd do differently
Read the reference images before building the report page. The first version was built from a written summary of the design references and came out as a 7,600-pixel essay — thirteen sections, ten screens of scroll — when every actual reference was a card grid with a headline-metric band. A day was spent learning that a doc's summary of an image loses exactly the thing that mattered.
And I would batch the evidence filter from day one. Scoring every item in a single model call worked until streaming delivered what we were paying for, at which point the call outgrew its own timeout and killed runs that had already spent money. The fix — small batches, bounded parallelism — was always the right architecture; it just wasn't the first one.
Worth naming what solo meant here: product, design, code and copy are mine, with AI-assisted engineering doing the accelerating and a set of third-party research APIs doing the reaching. The judgement calls — and the mistakes above — are all mine.



