Grok 1.5V
xAI · 2024-04-12
Grok 1.5V is xAI's first multimodal model, previewed in April 2024 as an extension of Grok 1.5 with the ability to process visual inputs including documents, diagrams, charts, and photographs alongside text. It was designed to demonstrate real-world spatial understanding and document-based question answering, marking xAI's initial entry into vision-language modeling. The model shared the core text capabilities of Grok 1.5 while adding image perception for multimodal tasks.
Benchmark scores
| Benchmark | Score |
|---|---|
| MathVista (mini) | 52.8 |