Can you tell that I am wearing smart glasses in this photo?
My new toy: RayNeo X3 Pro. Egocentric AI at your fingertips – or pupils? More comfortable to wear than one might expect – I genuinely forget whether I am wearing these or my usual sunglasses.
How well does this egocentric AI work? I was especially curious about the performance on visual inputs.
Well, when it works, it’s nothing short of magic. In-depth descriptions of scenes, places, settings. Correct descriptions of products. Speech understanding is superb. Very, very cool technology. The promise is clearly there.
But, in my day-to-day wear, in settings where I naturally wanted to use it, the AI failed a lot:
- 🫤 Did not answer questions, and there was no way to tell why: did it miss the voice command, lose connectivity, or was taking an unusually long time to reason?
- 🫤 Fabricated answers – spectacularly! All generative models do this to an extent, but what is wild about it in egocentric mode is that it confidently describes to you what you can clearly see with your own eyes is unambiguously incorrect – which feels different than when processing images on a screen. When evaluating a view of my sleek, lightweight road bike – “this is a BMX bike.” When asked to name people on a Zoom call, confidently gave four entirely made-up names. On many occasions, gave the same answers when asked to reconsider. When asked to summarize a presentation accompanied by a slide, invented details, like a student who is summarizing an article they have not read.
- 🫤 In suboptimal light conditions — perfectly workable for humans, but not ideal for AI, like in a dark room with a lamp placed on a table — was inaccurate to the point of being unusable.
Overall, latency and other constraints of this mode of operation give AI a distinct feel of being “so 2024.” The performance of cutting-edge, non-real-time VLMs, on conventional, non-egocentric image data is so much better than what is delivered in this regime.
Lessons? There is so much exciting work here for mobile systems, edge AI, HCI, and other research communities. So many important problems to solve, across so many dimensions.
- 📡 Edge AI continues to be important: the cloud-connectivity requirement is a major hassle in practice — and, conversely, improving connectivity in the open world would give this egocentric AI application a huge boost.
- 🤐 AI models must learn to know when not to provide an answer instead of delivering all results with the tone of confident self-assurance. Can humans ever learn to work with “assistants” who whisper confidently wrong judgments in their ear? The cognitive-load burden of dealing with this seems astronomical.
- 🛠️ The future is clearly agentic. Scene descriptions are fun to play with, but the real value will come when action follows perception — with human oversight, including the ability to override model judgment, for the foreseeable future.