Gemini 1.5 Flash-8B
Google DeepMind · 2024-10-03
Gemini 1.5 Flash-8B is Google's smallest Gemini API model; the exact stable gemini-1.5-flash-8b-001 snapshot became generally available on October 3, 2024 after experimental Flash-8B versions appeared in August and September. This multimodal transformer decoder is designed for high-throughput, low-latency uses such as chat, transcription, and long-context translation; the hosted model accepts text, images, audio, and video, returns text, and supports up to 1,048,576 input tokens. Google describes Flash-8B as a single-digit-billion-parameter model but does not disclose its exact total or active parameter count.