← Back to Blog

The Future of Voice and Vision: GPT-4o vs Google Project Astra

AI

May 2024 brought a seismic shift in how we interact with Artificial Intelligence. Within 24 hours of each other, OpenAI unveiled GPT-4o and Google showcased Project Astra, redefining the benchmark for multimodal AI.

GPT-4o: Omni-modal Real-time Interaction

GPT-4o ("o" for omni) processes audio, vision, and text in real-time, natively. The latency for audio responses has dropped to ~320 milliseconds—similar to human conversation. It can detect emotion, handle interruptions gracefully, and analyze live video feeds from a smartphone camera.

Google Project Astra: The Universal Assistant

Unveiled at Google I/O, Project Astra represents Google's vision for a universal AI agent. Astra demonstrated incredible continuous video understanding, where the agent could recall where objects were placed (like a user's glasses) and identify components of complex diagrams in real-time.

The Implications for Software Engineering

For studios like Akrevia, these advancements open entirely new paradigms for software interfaces:

  • Frictionless Field Tools: Imagine offline-first apps where field workers can simply point their camera and talk naturally to diagnose machinery, rather than filling out complex forms.
  • Accessibility: Real-time vision-to-voice changes how visually impaired users interact with software.

The race to ultra-low latency, multimodal AI is on, and the possibilities for custom software are boundless.