Back to Blog

From Project Astra to Gemini Spark: Inside Google's Universal AI Assistant (2026 Deep Dive)

Published ToolsBear Research Team
From Project Astra to Gemini Spark: Inside Google's Universal AI Assistant (2026 Deep Dive)

In May 2024, Google DeepMind showed a short video that quietly reset expectations for what an AI assistant is supposed to be. A researcher walked around an office holding a phone. She pointed the camera at a speaker and asked what part of it made sound; the assistant answered instantly. She asked where she had left her glasses — and the AI, having seen them on a desk a minute earlier, told her. The project was called Project Astra, and its stated ambition was blunt: build a universal AI assistant that can see, hear, remember and act in the real world.

Two years on, Astra is no longer a demo. Its capabilities have been folded, piece by piece, into products hundreds of millions of people use: Gemini Live's camera and screen sharing, persistent memory, the 24/7 Gemini Spark agent, the Daily Brief, and — shipping this fall — the first Gemini-powered smart glasses from Warby Parker and Gentle Monster. This article traces that path, explains what the assistant can genuinely do today versus what is still prototype, and looks at what "universal assistant" is going to mean for the way we use computers.

From Project Astra to Gemini Spark: inside Google's universal AI assistant

Executive Summary

  • Astra was a research prototype, not a product. Google describes it as the place where "breakthrough capabilities" are explored before graduating to Gemini Live, Search and new form factors like glasses.
  • The four capabilities that define it — see (camera/screen), hear (natural, interruptible audio), remember (across devices and sessions) and act (tools and interface control) — have all now shipped in some form.
  • The 2026 shift is from answering to doing. Gemini Spark runs multi-step, scheduled tasks across Docs, Sheets, Drive and the web for days or weeks, even when the app is closed. Daily Brief assembles your morning proactively.
  • Voice is the primary interface. Google reports 63% of Gemini users talk to it out loud, which is why the August 2026 Gemini Live upgrade centred on turning speech into action.
  • Glasses arrive in fall 2026 — audio-only first (cameras, mics, speakers, no display), display glasses later, plus XREAL's wired Android XR glasses. Each is a new "body" for the same assistant.
  • The endgame is a world model. Demis Hassabis framed the goal at I/O 2025: extend Gemini into a system that can simulate aspects of the world, plan, and take action on your behalf across any device.

1. What Project Astra Actually Was

The name matters less than the design brief. Google DeepMind set out to build an agent that could process a continuous stream of video and audio, respond with human-like latency, and hold context over time — "intuitively start conversations, and respond in the moment, without interrupting or time lag," as DeepMind's project page puts it. That is a very different engineering problem from a chatbot that waits for a typed prompt.

Three technical bets underpinned it:

  1. Native multimodality. Rather than bolting a vision model onto a text model, Gemini was trained from the ground up on text, images, audio and video, so it could reason about a camera feed directly.
  2. Streaming and memory. Astra encodes video frames and speech into a rolling timeline of events and caches it, which is what lets it answer "where did I leave my glasses?" — it has a short-term memory of what it saw.
  3. Native audio output. Instead of text-to-speech pasted on the end, the model generates speech directly, giving it natural intonation and the ability to be interrupted mid-sentence.
The four pillars of Project Astra: see, hear, remember, act
Astra's design brief in four verbs. Each has since shipped in a Gemini product.

Astra was, and still is, gated to a limited group of trusted testers. That is the point: it is DeepMind's sandbox. What matters to the rest of us is what graduates out of it.

2. The Two-Year Migration into Gemini Live

Google's own framing is that "some of the latest features in Gemini Live were first explored using Project Astra." The migration happened in stages.

Timeline from the Project Astra demo in May 2024 to Gemini smart glasses in fall 2026
From research demo to shipping hardware in roughly 30 months.

2024: Gemini Live launches

Three months after the Astra demo, Gemini Live arrived as a voice-first conversation mode inside the Gemini app: talk naturally, interrupt, change subject. It lacked vision at first, but it established the interaction model.

2025: Eyes and memory

Over 2025, camera and screen-sharing — the signature Astra abilities — rolled into Gemini Live so users could point a phone at the world or share a screen and talk about what was visible. Google also upgraded voice output to native audio, improved memory, and added computer control in testing. At I/O 2025, Hassabis published the roadmap explicitly: make Gemini a "world model" and turn the Gemini app into a universal assistant that "will perform everyday tasks for us, take care of our mundane admin and surface delightful new recommendations."

2026: From answering to acting

I/O 2026 (May 19) was the inflection point. Alongside Gemini 3.5 Flash — pitched as "frontier intelligence with action" and Google's strongest agentic model to date — the company introduced a redesigned Gemini app and two agents: Daily Brief, a personalised morning digest, and Gemini Spark, "a 24/7 personal AI agent designed to proactively manage tasks." Gemini Live's voice mode was merged directly into the main chat, so a conversation can move between typing and talking without losing context. A macOS app was announced with Spark able to operate on the local machine.

Then in August 2026 Google shipped the agentic upgrade to Gemini Live itself: voice commands that set up multi-step Spark tasks, hands-free Gmail triage (search, sort, delete by voice), a spoken Daily Brief, and memory that draws on past chats and connected app data.

3. Chatbot vs. Universal Assistant: What Actually Changed

Comparison table of a classic chatbot versus a universal AI assistant

It is tempting to see this as feature creep on a chatbot. It is more useful to see it as a change in where the assistant lives in your day:

  • Input moves from a typed prompt to whatever is in front of you — your camera, your screen, your voice.
  • Memory moves from a single conversation to a persistent, cross-device profile: dietary preferences, family dates, past decisions, your app data (with permission).
  • Initiative flips. Daily Brief does not wait to be asked; it shows up with what you need to know.
  • Action replaces answers. "Sort my inbox" or "book my usual coffee order" ends with something done, not a paragraph of instructions.
  • Time horizon stretches from seconds to weeks. Spark handles "long-running, scheduled jobs across days or weeks, even when you aren't actively using the app."
  • Surface expands from one app to phone, desktop and glasses, with the conversation following you between them.

4. The Architecture Behind the Assistant

Layered architecture: devices, Gemini Live interface, agents, models, tools
A simplified view of how the pieces layer. The assistant is the interface; the agents do the work; the models supply the intelligence.

Think of five layers:

  1. Devices — phone, the new macOS app, and from this fall, glasses. Cross-device memory means you can start on one and continue on another.
  2. Gemini Live — the unified voice-plus-text conversation surface. The re-engineered microphone lets you "talk through a complex idea at your own pace without getting cut off."
  3. Agents — Spark (long-running tasks) and Daily Brief (proactive summaries). These are what make the assistant "do" rather than "say."
  4. Models — Gemini 3.5 Flash today (3.5 Pro announced to follow), plus Gemini Omni for turning prompts into video. Speed matters here: Google claims 3.5 Flash outputs tokens roughly four times faster than other frontier models, which is what makes real-time voice feel real-time.
  5. Tools — Search, Gmail, Calendar, Maps, Docs/Sheets/Drive and, increasingly, third-party apps that the assistant can operate on your behalf.

5. Gemini Spark: The Agent That Keeps Working

Flow of a Gemini Spark task from voice request to completed report

Spark is the clearest expression of the Astra thesis. You describe a goal — by voice, if you like — and Spark plans the steps, works across your Google apps and the web, and keeps the job running on a schedule. Google's examples lean practical: monitor something weekly and compile a sheet, turn messy notes into a structured document, keep a project tracker updated.

What makes it different from earlier "agents" is persistence. The task survives you closing the app. It remembers the goal. It can run over weeks. That is the difference between a tool you operate and a colleague you delegate to — and it raises the same question you would ask of a new colleague: how much do you trust it, and how do you check its work? Google's answer, repeated throughout its announcements, is "all under your direction" — Spark acts within the permissions and connections you grant.

6. Smart Glasses: A New Body for the Assistant

Three types of Gemini eyewear: audio glasses, display glasses and wired XR glasses

Astra was always shown running on prototype glasses, because glasses are where "see what I see" stops being a phone gimmick. At I/O 2026 Google confirmed the roadmap for Android XR eyewear built with Samsung and Qualcomm:

  • Audio glasses — shipping fall 2026. Cameras, microphones and speakers but no display. Say "Hey Google" or tap the frame; ask about what you're looking at, get turn-by-turn audio directions from Maps, send texts, take photos, and trigger agentic actions in phone apps (Google's demo ordered "my usual" from a coffee shop through DoorDash by voice). Frames come from Warby Parker and Gentle Monster and pair with both Android and iOS.
  • Display glasses — later. Add an in-lens private display for navigation arrows and live translation captions.
  • Wired XR glasses — XREAL Aura, fall 2026. A different category: optical see-through with a 70° field of view, running Android XR on Snapdragon, deep Gemini integration, tethered to a battery pack. More "portable spatial workspace" than "all-day assistant."

The strategic logic is straightforward. Meta's Ray-Ban glasses proved people will wear AI eyewear if it looks normal. Google's bet is that Gemini's depth of integration with Maps, Gmail, Android apps and agentic actions is the differentiator — and that the same assistant that lives in your phone should simply be present in your glasses too.

7. Memory, Privacy and the Trust Problem

What the assistant can remember and the controls users have

An assistant that remembers what you said last month, reads your inbox and sees through your glasses is only useful if you trust it — and only trustworthy if you can see and control what it holds. Google's public position is that memory is opt-in and reviewable, app connections are granted individually, and personalisation ("Personal Intelligence," in Google's August 2026 phrasing) draws on your data only with permission.

Users should still think carefully about three things: what the camera captures in public when glasses are worn; which app connections an agent has been given (an agent that can delete emails by voice is powerful in both directions); and how to audit what a long-running Spark task did while you weren't watching. None of this is unique to Google — every company building agents faces it — but the more capable the assistant, the more these controls matter.

8. The World-Model Endgame

Cycle of perceive, simulate, plan and act describing Gemini as a world model

Why does Google keep using the phrase "world model"? Because a truly universal assistant has to do more than retrieve and summarise — it has to anticipate. Hassabis's I/O 2025 essay described extending Gemini into a system that can "make plans and imagine new experiences by understanding and simulating aspects of the world, just as the brain does." DeepMind's Genie 3, which generates interactive, physically consistent environments in real time, is the research arm of that idea (we cover it in depth in our world models explainer).

For the assistant, the practical payoff is planning: an AI that can simulate "if I reorder these tasks, what happens to the deadline?" or "if I take this route, will I make the train?" is one that can act with judgement, not just follow instructions.

9. Where It Sits Against the Competition

Positioning grid of AI assistants by proactivity and device reach
Editorial positioning, not a benchmark.

OpenAI is pursuing the same destination from the other direction — its GPT-5.6 models and Agents API emphasise long-running professional agents and computer use, and its consumer push centres on ChatGPT rather than a hardware ecosystem (for now). Meta owns the glasses beachhead with Ray-Ban Meta and its Muse Spark model, but lacks Google's depth in maps, mail and productivity apps. Apple's Siri remains largely reactive and device-bound. Google's distinctive claim is reach: one assistant, one memory, across phone, desktop, Search and glasses, with agents that act inside the apps you already use.

10. What This Means for You

Six everyday use cases for the Gemini universal assistant

Practically, here is what you can do today and what to expect next:

  • Today (Gemini app): talk and type in one thread; point your camera or share your screen and ask questions; get a spoken Daily Brief; triage Gmail by voice; set up Spark tasks that run over days; let Gemini remember preferences across sessions.
  • Fall 2026: audio glasses that bring the same assistant to your face — directions, photos, messages, app actions by voice — plus XREAL Aura for a wearable Android XR workspace.
  • Coming: Gemini 3.5 Pro, display glasses, deeper computer control on desktop via the macOS app, and regional voice dialects.
63 percent of Gemini users talk to it out loud

The habit to build now is delegation. Astra's lineage is about an assistant that does things, and the users who benefit most will be the ones who learn to hand over bounded, checkable tasks — "compile this weekly," "keep this tracker updated," "clear the newsletters" — and then verify the results, exactly as you would with a new assistant of the human kind.

Key takeaways

  • Project Astra is the research lab; Gemini Live, Spark, Daily Brief and the glasses are what shipped from it.
  • The 2026 turn is from answering questions to executing multi-step tasks over time, driven mostly by voice.
  • Glasses (audio first, fall 2026) make "see what I see" the default rather than a phone trick.
  • Memory and app permissions are the trust surface — review them deliberately.
  • The long game is a world model that can simulate and plan, not just retrieve.

Frequently Asked Questions

Is Project Astra something I can download?

No. Astra is a research prototype available to a limited group of trusted testers via a waitlist. Its capabilities reach the public through the Gemini app (Gemini Live, Spark, Daily Brief) and, from fall 2026, Gemini-powered glasses.

What is the difference between Gemini Live and Gemini Spark?

Gemini Live is the conversational interface — voice and text, camera and screen sharing. Spark is an agent you can launch from that conversation to carry out multi-step, scheduled tasks across your apps and the web, continuing after you close the app.

When do the Google smart glasses come out and do they have a screen?

The first Gemini glasses from Warby Parker and Gentle Monster ship in fall 2026 and are audio-only: cameras, microphones and speakers, no display. Display glasses with an in-lens screen are planned later. XREAL Aura, a wired Android XR glasses device with a display, is also due this fall.

Does Gemini remember everything I say?

Gemini can remember key details across sessions and use connected app data for personalisation, but memory is designed to be reviewable and controllable, and app connections are granted individually. Check the memory and connected-apps settings and prune what you don't want retained.

What is a "world model" and why does Google keep mentioning it?

A world model is an AI that can predict how the world (or a task) will evolve in response to actions — it simulates rather than just retrieves. Google's stated goal is to extend Gemini into one so the assistant can plan ahead and act with judgement.


Sources: Google DeepMind Project Astra page; Google I/O 2025 keynote essay "Our vision for building a universal AI assistant" (Demis Hassabis, May 2025); Google blog "The Gemini app becomes more agentic" and "Gemini 3.5: frontier intelligence with action" (May 19, 2026); "Turn your voice into action with new productivity features in Gemini Live" (Aug 26, 2026); Google Android XR announcements at I/O 2026 and The Android Show XR edition; XREAL Aura launch release (June 2026); UploadVR reporting on Gemini smart glasses (June 2026). Product claims are as publicly stated by the companies.

TO

ToolsBear Research Team

Research & Editorial

Written by the ToolsBear Research Team team. We test tools, study market trends, and turn complex topics into clear, actionable guides you can use for your next project.