The speech-to-text model provides real-time transcription alongside screen context to summarize local files, conduct research, and generate images.
macOS Gemini users can now use real-time speech and active screen context to process local documents and desktop tasks.