Spokes.wiki Search About
Article source ↗ source url updated Thu Aug 13 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Have LLMs Finally Mastered Geolocation?

Bellingcat (Foeke Postma and Nathan Patin, 6 June 2025) ran 20 models from five developers (OpenAI, Google, Anthropic, Mistral and xAI) over 25 unpublished travel photographs from every continent, stripped of metadata and chosen to vary in difficulty. Each answer was scored 0–10, where 10 meant “accurate and specific identification, such as a neighbourhood, trail, or landmark,” and the whole field was benchmarked against Google Lens’s top ten visual matches.

This is the first source in this spoke where a practicing investigative organisation measures an AI at an OSINT task it does itself, against ground truth it holds. llm-osint-reliability-study scored models on analysing evidence handed to them; this scores them on a specific craft skill with a right answer.

What it found

Only three models beat Google Lens overall: ChatGPT o3, o4-mini and o4-mini-high. Below them came Grok (xAI’s DeeperSearch strongest of that family), then Gemini, then Claude and Mistral at the bottom. The unpublished-photo design matters: a model cannot have seen these images, so a correct answer has to come from reading the picture rather than recalling it.

The failure modes are the useful part

  • Every model hallucinated at some point. “All models produced entirely wrong answers at times.”
  • The reasoning modes made things worse. “Deep research” and “extended thinking” settings typically scored below the same models’ standard modes.
  • Account history leaked into answers. One ChatGPT instance drew on the tester’s previous conversations, so the result was not a function of the image alone.
  • Confidence tracked both ways. ChatGPT’s assertiveness produced better answers and more hallucinations.
  • Temporary and seasonal features (installations that were not there when the reference imagery was taken) defeated the models.
  • Video comprehension remains “limited.”

The authors’ conclusion is deliberately unexcited: “LLMs are no silver bullet. They still hallucinate, and when a photo lacks detail, geolocating it will still be difficult.” What they grant is narrower: models “help researchers to spot the details that Google Lens or they themselves might miss,” strongest in urban scenes where multilingual signage and small visual cues are in play.

Why it matters here

The pattern this spoke keeps recording is that the AI-OSINT claims are asserted by the tools’ own authors. Here the assessor is the institution that would use the capability, it published the losers along with the winners, and it named a defect, the account-history leak, that no benchmark scoring a fixed prompt set would ever surface.

It also cuts against the assumption behind the spoke’s autonomous tools. llm-osint and kallisto-osinter chain model calls into a pipeline; if the reasoning modes score worse than plain prompting on the one task here with a ground truth, more scaffolding is not obviously more accuracy. That is a measured result pointing the opposite way from the tools’ design.

What it still does not answer: geolocation is analysis of an image someone already has. Nothing here measures whether an autonomous collector finds the right material to begin with, which is this spoke’s standing gap.

Connections

bellingcat · geolocation-chronolocation · llm-osint-reliability-study · ai-osint · llm-osint · kallisto-osinter · glan-bellingcat-methodology