Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How do you think humans experience desktop interfaces? “Basically just OCR'ing screenshots” is exactly what humans do.


It's not the same thing. For example, given a GUI with a titlebar, title, subtitle, text, and buttons, a human can instantly understand spatially the relationship between these items. But a naive OCR of such a GUI would be a flat stream of text that loses a ton of information.


But that’s not how models handle images either. They spatially segment and reason about title bars, placement, etc.


I was also under the impression modern AI agents have moved on from just OCR'ing screenshots to leveraging native vision model capabilities.


They do. They all use ViTs and have for quite a while.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: