Spokes.wiki Search About
Scholarly Article source ↗ source url updated Fri Aug 07 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Evaluating the Reliability of LLMs in OSINT Investigations: A Friend or Foe?

Peer-reviewed conference paper, ECCWS 2026 (25th European Conference on Cyber Warfare & Security, June 2026), by Errol Baloyi, Nokuthaba Siphambili, Ntomfuthi Ntshangase, Mpho Letshwenyo, Fhatuwani Makharamedzha, Rendani Mmbodi and Ndabezinhle Hlongwane of South Africa’s Council for Scientific and Industrial Research (CSIR). DOI 10.34190/eccws.25.1.4818.

The spoke’s second T1 and — more to the point — its first measurement of anything. Every accuracy claim in this wiki up to now has been self-asserted by the tool that makes it.

What they did

Built a fictional organization, the “CtrlZ Society,” with a fabricated online footprint, then fed eight LLMs synthetic evidence of it across the shapes a real investigation meets: social-media posts, Reddit threads, Pastebin documents, GitHub repositories, blog posts, Telegram broadcasts, WHOIS records and news articles. Models tested: GPT, Claude 3, Gemini, Meta, DeepSeek, Qwen 2.5, Mistral Large and Grok.

Scoring ran on six criteria: accuracy · timeline reconstruction · account attribution · evidence traceability · hallucination susceptibility · handling of ambiguous or incomplete data.

A real target would make ground truth unknowable and the ethics unworkable. Inventing the organization means the authors know every true fact and can score a model against it, which is what lets them tell a correct retrieval from a confident invention.

What they found

  • Claude 3, GPT and Qwen 2.5 gave “strong analytical performance and reliable synthesis of investigative outputs.”
  • Gemini and DeepSeek performed less well.
  • Meta and others showed “forced narrative construction when prompted adversarially” — pushed toward a conclusion, the model builds the story rather than declining. For OSINT this is the worst available failure: an investigator who leans on a hypothesis gets it confirmed.
  • All eight were nonetheless useful for structuring and summarizing complex investigative material.
  • The authors’ recommendations are human oversight, multi-model validation, and verification protocols — not “don’t use them.”

What it settles here

Two of this spoke’s open questions move.

“Where does the human stay in the loop?” now has an answer with evidence behind it rather than a design preference. The paper’s case for oversight is not that autonomy is distasteful — it is that adversarial or leading prompting produces fabricated narrative, so the human is the control for a measured failure mode. That is a stronger argument than kafsiem‘s auditable-edge design makes on its own, and it points the same way as the berkeley-protocol‘s professional duties.

“No measurement of the AI-automation claims” (growth edge #2) is answered in part. The measurement exists, it is peer-reviewed, and it is unflattering in a specific way rather than generally. Note the scope limit: this scores LLMs given evidence, not the autonomous gatherers like llm-osint or kallisto-osinter, whose collection accuracy is still unmeasured. The analysis half now has a number; the collection half does not.

Caveats

  • Model names are coarse. “GPT,” “Claude 3,” “Meta” and “Grok” name families, not versions, and by ECCWS 2026 several are well behind current releases. Treat the per-model ranking as dated; treat the failure mode — forced narrative under adversarial prompting — as the durable finding.
  • A synthetic target is not a real one. It buys ground truth and costs realism: no stale data, no deliberate counter-OSINT, no ambiguity the authors did not author.
  • This summary rests on the paper’s published abstract and record, not the full text.
  • ai-osint — the claims this measures
  • llm-osint — an autonomous profiler whose accuracy this does not cover
  • kafsiem — analyst-in-the-loop as design; this is analyst-in-the-loop as finding
  • berkeley-protocol — the duty-side argument for the same conclusion
  • synthesis