A developer ran a controlled benchmark to test whether a widely-praised 'magic' GeoGuessr prompt actually improved OpenAI o3's geolocation accuracy. Using 200 images and comparing the elaborate prompt against a basic one, the results showed the default prompt consistently performed better. The finding illustrates how easy it is to fool yourself about prompt engineering effectiveness when a model is already capable — iterating with the model and asking it to evaluate its own improvements produces unreliable feedback. The benchmark also revealed that newer models (gpt-5.4 and gpt-5.5) do not match o3's geolocation capabilities, suggesting that specific capability was not carried forward.

6m read timeFrom seangoedecke.com
Post cover image
19.4K Impressions