RefCOCO
Multimodal · 2016-07-31
RefCOCO is a referring expression comprehension dataset for localizing objects in images from natural-language expressions. Scores are usually reported as localization accuracy on validation and test splits.
Top models (higher is better)
| Model | Score |
|---|---|
| Interfaze Beta | 82.1 |
| Gemini 3.5 Flash | 80.9 |
| Gemini 3 Flash Preview | 75.2 |
| Sonnet 5 | 69.2 |
| GPT-5.4 Mini | 67.0 |
| Grok 4.3 | 25.0 |