
Oxford Psychology Series: Active Vision: The Psychology of Looking and Seeing. Findlay & Gilchrist, 2003
I have used this line so many times myself: smart glasses see what you see and hear what you hear. It is well known in the world of egocentric AI and AR; it appeared verbatim on Meta’s quarterly earnings call back in 2023 (“glasses are the ideal form factor for an AI device because they enable your AI assistant to see what you see and hear what you hear“). To a first-order approximation, this statement is correct: human sensors (eyes; ears) and machine sensors (cameras; microphones) are positioned such that largely the same signals reach them.
But it is also misleading: smart glasses may see what you see, but they don’t look the way you look; they may hear what you hear, but they don’t listen the way you listen. Looking and listening are dynamic, active processes, steered to get the information needed to interpret the situation. When an object is partially occluded, you tilt your head slightly to look over the obstacle. When glare prevents you from reading what you need to read, you angle what you are reading away from it. You move closer to or farther from what interests you; your eyes focus where needed; and both vision and hearing are continuously supplemented by information from other senses.
Even if AI became good enough to evaluate captured video and audio as well as humans do, which it currently is not, it would still not reach human-level performance in day-to-day use if it remained unable to actively seek the information it was missing. Even if each and every ML-model issue were fully addressed, egocentric AI would still face a fundamental limitation: it lacks the agency to collect, on the fly, the data it needs to improve its judgment.