- Role
- Designed and built by me in Python
- Model
- A local vision-language model (9B) served through LM Studio
- Speed
- One look every one to two seconds
- Result
- Three working versions: observe, act, and learn
Summary
An agentic AI prototype for computer use. A local vision-language model (multimodal AI) reads the screen and returns the state and the next action as structured JSON, which is the core of vision-based RPA for systems that have no API. A second version reads the numbers with OCR and learns what works with reinforcement learning (Q-learning). It all runs on local LLM inference, so the screen contents stay on the machine, and acting is gated: observe-only by default, an explicit switch, a pause between actions and a kill switch (human-in-the-loop).
Problem
A lot of business software has no API. Old ERP screens, public portals, remote desktops and vendor tools can only be used the way a person uses them: by looking at the screen and pressing keys. That makes them the hardest part of any automation project. I wanted to know how far a small model running on a normal PC can get by simply looking at the screen, and what it takes to let it act safely.
What I built
Three versions, each testing one idea, on a browser-based turn-based game. A game is a good test bench: the screen changes all the time, the rules are clear, and a wrong move costs nothing.
The observer looks only. A first pass sends one full screenshot to a local vision-language model, which finds the parts of the screen that matter and returns their coordinates. After that, every look crops just those regions into one small image, and the model reports the state as JSON, about once a second.
The agent also acts. It finds the game area, and on every look the model returns the state and one proposed action, a key press or a click, as JSON. Clicks are mapped back from the cropped image to real screen coordinates. Acting is off when it starts: it runs in observer mode until I switch execution on, keeps a minimum pause between actions, and stops at once if the mouse is pushed into a corner.
The learner drops the language model. It attaches to the browser over Chrome DevTools, reads the numbers on screen with OCR, turns them into a simple state, and learns which key to press with Q-learning, rewarded for damage dealt and penalised for damage taken. What it learns is saved between runs.
The thinking behind it
The design follows the same rule as my other work: use the model where understanding is needed, and plain rules where the numbers have to be right. The vision model is good at the open question of what is on the screen and what to do next. Reading an exact number off a bar is better done with OCR and a pattern check, which is cheap and either matches or fails.
Looking at the whole screen once and then only at the parts that matter keeps each look small, which is what makes a 9B model on a home PC fast enough to keep up. Comparing the model's decisions with a learned policy shows the trade-off: the model can act sensibly from the first second but is slow, while reinforcement learning is fast and cheap per step but only as good as the state it is given.
Where it stands
All three versions run end to end on my own PC. The observer reports the screen state about once a second, the agent plays through screens on its own once switched on, and the learner builds up and saves its own table of what works.
Where it is useful, and what is next
The same loop (look, propose, act only through a gate) fits work where the screen is the only interface:
Data entry into legacy systems, where the agent reads a document or a screen and fills in another system that has no API.
Monitoring, using the observer alone: it watches a dashboard, a queue or a production screen and raises a flag when something changes, without ever touching it.
Testing, where it walks through an application the way a user would and reports what it sees.
For that, the next steps are a log of every look and every action for audit, a fixed list of allowed actions per task, a confirmation step before anything that writes data, and running the observer on a real business screen before letting the agent act on one.
Stack: Python, a local vision-language model through LM Studio, mss, PyAutoGUI, Playwright, Tesseract OCR.
Code: On request.