GPT-5.6 Sol Leads OpenAI Vision Tests

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- GPT-5.6 Sol scored 46.2 mAP@50 in Roboflow’s object-detection test, up from GPT-5.5’s 13.8; Terra and Luna followed at 44.7 and 43.3.
- GPT-5.6 Sol reached 73.0% in object counting, up from GPT-5.5’s 64.9%, and correctly handled overlapping metal brackets and count rules restricted to selected scoring zones.
- GPT-5.6 Sol recorded a 90.7% mean OCR similarity score, 0.5 points below GPT-5.5’s 91.2%, while its 82.5% text-extraction score trailed the prior model’s 87.6%.
- GPT-5.6 models returned their best detection results with absolute XYXY pixel coordinates; the wrong format cut performance by about 15 mAP points, while Gemini 3.5 Flash favored normalized 0–1000 YXYX coordinates.
- OpenAI confirmed that Sol becomes less stable on images around 2,000 by 2,000 pixels or larger at lower reasoning effort; higher reasoning improved stability but increased token use, latency, and cost, and resizing or cropping was identified as the practical workaround.
- GPT-5.6 Luna finished in slightly over 5 seconds per image, versus about 6 seconds for Terra and close to 10 seconds for Sol, and cost less than 0.5 cents per image.
- Gemini 3.5 Flash led the benchmark’s detection and counting tests and cost about 0.8 cents per image, compared with roughly 2.5 cents for Sol, making it a strong option for data-intensive workloads.
Why it matters: Developers building agents and document workflows gain a much stronger OpenAI vision option: Sol’s detection rose to 46.2 mAP@50 and counting to 73.0%, but teams processing large image batches must weigh its 2.5-cent cost and instability above roughly 2,000-by-2,000 pixels against cheaper rivals such as Gemini 3.5 Flash.
Ask SkimNews




