Eyes - The Hybrid Desktop Agent (Claude+qwen)
About this tool
Eyes is a native Windows 11 agent that takes an instruction in plain language and carries it out in your real applications — Word, Excel, Outlook, Teams, File Explorer, VS Code, your browser — by seeing the screen, reading the UI tree, and acting like a person would. It does not record macros and it does not rely on brittle selectors that break when a button moves. You describe the outcome; it works out the steps. It runs entirely on your machine. Your files never leave it. What it does Executes multi-step tasks across applications. Read these invoices, extract the totals, build a spreadsheet, attach it to an email. One instruction, many apps, one result. Operates real Windows software through UI Automation, COM, OCR and visual grounding — with a learned click memory that gets faster at what you do often. Works with documents: reads and writes Word, Excel and PDF, indexes local document sets, and answers questions from them. Talks and listens. Chat window, floating overlay, and voice. Plans, re-plans, and repairs itself when an application isn't where it expected it to be. Ships with 116 tools available to the model and prebuilt skill packs for Word, Excel, Outlook, Teams, Explorer, VS Code, the terminal and LibreOffice. What makes it different Most desktop agents will tell you a task is done. Eyes will tell you how it knows. Verification is built into the execution path, not bolted on. Before a task is closed, a postcondition gate checks that the promised artifact actually exists on disk, with the promised name, created during this task — not a similar file from yesterday. If the file isn't there, the task is not "done". Three-valued outcomes. A task ends as achieved, failed, or unknown. An agent that cannot see whether it succeeded says so instead of assuming the best case. This sounds like a small thing. It is the difference between an automation you can audit and one you have to babysit. Gates that fail closed. When Eyes is about to do something destructive and cannot confirm what window is in the foreground, it stops. When an approval channel is unavailable, it refuses rather than proceeding. Not being able to see is treated as a reason to halt, never as permission to continue. A receipt for every run. Each task produces a record of what was attempted, which tools were called, which safety gates fired, what was verified and against what evidence, how long it took and what it cost — plus a deterministic self-assessment that flags repeated calls, guard collisions, dead time and rejected deliveries. Privacy that is enforced, not promised. Sensitive data is pseudonymised before anything crosses the network, a local-model mode is available for work that must not leave the machine at all, and automated checks verify that no screenshots, credentials or session tokens are included in what is sent. Cost routing that measures instead of guessing. The model router computes the actual break-even before downgrading to a cheaper model, using live token prices rather than hardcoded constants, and never pays the switching cost on a task that is already lost. Engineering This is not a prototype with a demo video. ~191,000 lines of Python, single author. 7,599 passing tests across 367 test files. 54 closure checks that assert architectural invariants — one door for file writes, no internal call that bypasses a safety gate, no silent exception growth, no promise in a prompt that the code doesn't keep. 53 of those 54 checks are verified by mutation testing: each one declares how to break it, and a tool confirms the check actually goes red when the thing it guards is broken. A check that cannot fail is not a check. Platform provenance is tracked: the suite refuses to claim "verified" unless it has been run on Windows with exactly this content, fingerprinted by hash. A green result from another operating system does not count. Cost Bring your own key. No subscription, no hosting, no per-seat fee. Measured on real runs: ~$0.25 per completed task (range $0.15–$0.40 depending on task length and how much visual context is needed). Normal use of around 15 tasks a working day comes to roughly €70/month in model spend. Light use is nearer €25; heavy use around €200. A local-model mode brings that to zero in exchange for capability. Prompt caching does most of the work here: the first call in a task carries the cost, and subsequent calls are an order of magnitude cheaper — so one long task is proportionally cheaper than five short ones. What this is not, as of today Included deliberately, because you will find this out anyway and I would rather you heard it from me. No third-party benchmark score yet. A run on Microsoft's Windows Agent Arena is scheduled; the number will be published with its method, its confidence interval, and a control arm running a generic agent on the same 154 tasks and the same machine. Until that exists, treat any capability claim here as the author's, not an independent result. No installed user base. This has been built and hardened by one person, not validated by a customer base. One known failing test in the suite, diagnosed: an integration test whose outcome depends on state left behind by previous runs. The diagnosis is written down; the fix is not yet in. Windows only. By design — the value is in the depth of the Windows integration, and porting it away would throw that out. Licence Available as a pilot licence for evaluation in your own environment. Broader licensing and IP transfer are discussed case by case. The author is available for integration and transition work.
What you will receive
- ✓Source code provided (private repo access or ZIP)
- ✓All dependencies and environment variables documented (.env templates)
- ✓Database schema / data model documented
- ✓Setup & self-hosting guide handed over
- ✓Walkthrough call completed (min. 30 min)
- ✓License terms confirmed (non-exclusive; brand & trademark not included)
Demo video
Charged in EUR