Building a Selector-Free Autonomous Desktop RPA Engine with Python & Vision LLMs
Legacy enterprise software—on-premise SAP GUI installations, AS/400 terminal emulators, desktop-only accounting suites, and specialized medical record software—rarely provides accessible REST APIs. For the past two decades, the standard remedy has been traditional RPA (Robotic Process Automation) platforms like UiPath, Automation 360, or Power Automate Desktop.
In practice, traditional RPA is notoriously brittle. It relies on accessibility element trees, Win32 handle introspection, or static DOM selectors. When an update alters an internal UI hierarchy, a display scale shifts from 100% to 125%, or dynamic IDs change, automation jobs break.
This guide demonstrates how to architect a selector-free, OS-agnostic RPA engine in Python that leverages Multimodal Vision LLMs (such as Anthropic Claude 3.5 Sonnet / Computer Use) to perceive screen buffers and drive low-level OS input events natively across Windows, macOS, and Linux.
1. The Bottleneck: Why Traditional RPA Breaks
Traditional RPA frameworks rely on accessibility trees (Microsoft UI Automation, Apple Accessibility API, AT-SPI). This architecture incurs three major vulnerabilities:
- Fragile Tree Traversal: Virtualized grids (like SAP or legacy WPF apps) render contents via hardware-accelerated canvases without exposing child accessibility nodes.
- Resolution & DPI Vulnerability: Pixel-offset OCR fallbacks shatter when remote desktop windows resize, rendering absolute coordinates invalid.
- Licensing Costs: Commercial enterprise runners charge $5,000 to $15,000 per unattended bot license annually, locking infrastructure into closed ecosystems.
The Vision-Grounding Alternative
Instead of querying internal OS handle structures, a Vision-Driven RPA engine operates identically to a human engineer:
- Capture the raw desktop framebuffer.
- Pass the downscaled/encoded buffer into a multimodal reasoning model alongside an action objective.
- Extract normalized visual coordinate predictions (
x_pct, y_pct) alongside an action token (click, double_click, type, key_press, scroll).
- Map those coordinates to native hardware events.
- Verify the state delta across consecutive frames with a self-healing loop.
+--------------------+ +-------------------------+ +-------------------------+
| Desktop Framebuffer| ---> | Vision Coordinate Model | ---> | Action Dispatcher |
| (X11 / Quartz / OS)| | (Claude / Qwen-VL) | | (pyautogui / pynput) |
+--------------------+ +-------------------------+ +-------------------------+
^ |
| +-------------------------+ |
+---------------- | Visual Verification Loop|<------------------+
| (Delta / SSIM Check) |
+-------------------------+
2. System Architecture
The engine consists of four core decoupling boundaries:
- Display Capture Pipeline: Intercepts screens via fast C-bindings (
mss) and handles multi-monitor resolution normalization.
- Coordinate Normalizer: Converts screen coordinates to a standardized
1000x1000 grid to avoid prompt overhead, then re-scales back to client coordinates.
- Action Execution Engine: Dispatches hardware inputs using
pyautogui or virtual input device drivers (/dev/uinput on Linux).
- Audit & Self-Healing Loop: Records state deltas using Structural Similarity Index (SSIM) and saves structured JSONL execution logs.
3. Implementation: Code & Logic
Step 1: Screen Capture and Coordinate Rescaling
We capture the screen buffer with mss and rescale coordinates between raw monitor resolution and normalized model space (0–1000).
import mss
from PIL import Image
import io
class ScreenCapture:
def __init__(self, monitor_index: int = 1):
self.sct = mss.mss()
self.monitor = self.sct.monitors[monitor_index]
self.width = self.monitor["width"]
self.height = self.monitor["height"]
def capture_frame(self, target_width: int = 1280) -> tuple[bytes, float]:
sct_img = self.sct.grab(self.monitor)
raw_image = Image.frombytes("RGB", sct_img.size, sct_img.bgra, "raw", "BGRX")
# Maintain aspect ratio for LLM visual processing
aspect_ratio = self.height / self.width
target_height = int(target_width * aspect_ratio)
resized_image = raw_image.resize((target_width, target_height), Image.Resampling.LANCZOS)
buffer = io.BytesIO()
resized_image.save(buffer, format="PNG", optimize=True)
# Scale factor relative to native resolution
scale_factor = target_width / self.width
return buffer.getvalue(), scale_factor
def denormalize_coordinates(self, norm_x: int, norm_y: int, grid_size: int = 1000) -> tuple[int, int]:
"""Translates normalized grid coordinates (0-1000) to actual monitor pixels."""
real_x = self.monitor["left"] + int((norm_x / grid_size) * self.width)
real_y = self.monitor["top"] + int((norm_y / grid_size) * self.height)
return real_x, real_y
We enforce structured JSON tool responses so the LLM outputs unambiguous actions and coordinates.
TOOL_DEFINITION = {
"name": "desktop_action",
"description": "Execute an input action on the desktop target UI.",
"input_schema": {
"type": "object",
"properties": {
"action": {
"type": "string",
"enum": ["click", "double_click", "right_click", "type", "press_key", "hotkey", "wait", "terminate"]
},
"coordinates": {
"type": "array",
"items": {"type": "integer"},
"description": "[x, y] coordinates mapped to a 0-1000 grid. Required for click events."
},
"text": {
"type": "string",
"description": "Text string to type into input fields."
},
"key": {
"type": "string",
"description": "Key or key combination to trigger (e.g., 'enter', 'tab', 'ctrl+s')."
},
"reasoning": {
"type": "string",
"description": "Step-by-step rationale for why this UI element was selected."
}
},
"required": ["action", "reasoning"]
}
}
Step 3: Hardware Action Execution & State Verification
The dispatcher executes hardware events and captures state confirmation immediately afterward.
import pyautogui
import time
import json
from datetime import datetime
pyautogui.FAILSAFE = True
pyautogui.PAUSE = 0.05
class ActionDispatcher:
def __init__(self, capture: ScreenCapture, log_path: str = "audit_log.jsonl"):
self.capture = capture
self.log_path = log_path
def dispatch(self, action_payload: dict) -> bool:
action = action_payload.get("action")
coords = action_payload.get("coordinates")
text = action_payload.get("text")
key = action_payload.get("key")
real_x, real_y = None, None
if coords and len(coords) == 2:
real_x, real_y = self.capture.denormalize_coordinates(coords[0], coords[1])
# Execute action
if action in ["click", "double_click", "right_click"] and real_x and real_y:
pyautogui.moveTo(real_x, real_y, duration=0.2)
if action == "click":
pyautogui.click()
elif action == "double_click":
pyautogui.doubleClick()
elif action == "right_click":
pyautogui.rightClick()
elif action == "type" and text:
pyautogui.write(text, interval=0.02)
elif action == "press_key" and key:
pyautogui.press(key)
elif action == "hotkey" and key:
keys = [k.strip() for k in key.split("+")]
pyautogui.hotkey(*keys)
elif action == "wait":
time.sleep(2.0)
elif action == "terminate":
return False
self._log_audit(action_payload, (real_x, real_y))
return True
def _log_audit(self, payload: dict, absolute_coords: tuple):
log_entry = {
"timestamp": datetime.utcnow().isoformat(),
"payload": payload,
"absolute_coords": absolute_coords
}
with open(self.log_path, "a") as f:
f.write(json.dumps(log_entry) + "\n")
Step 4: The Closed-Loop Execution Loop
This loop orchestrates the screenshot capture, calls the Anthropic API, executes the action, and checks whether the visual state changed.
import base64
from anthropic import Anthropic
def run_autonomous_task(task_instruction: str, max_steps: int = 25):
client = Anthropic()
capture = ScreenCapture(monitor_index=1)
dispatcher = ActionDispatcher(capture=capture)
messages = []
for step in range(max_steps):
print(f"[*] Executing step {step + 1}/{max_steps}")
img_bytes, _ = capture.capture_frame(target_width=1280)
base64_img = base64.b64encode(img_bytes).decode("utf-8")
step_prompt = (
f"Current task: {task_instruction}\n"
f"Analyze the provided desktop screenshot. Identify the target element "
f"and invoke the desktop_action tool. If the task is finished, invoke terminate."
)
messages.append({
"role": "user",
"content": [
{"type": "text", "text": step_prompt},
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": base64_img}}
]
})
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=1024,
tools=[TOOL_DEFINITION],
tool_choice={"type": "tool", "name": "desktop_action"},
messages=messages
)
tool_call = next(c for c in response.content if c.type == "tool_use")
action_data = tool_call.input
# Append assistant turn to prevent context drift
messages.append({"role": "assistant", "content": response.content})
continue_running = dispatcher.dispatch(action_data)
if not continue_running or action_data.get("action") == "terminate":
print("[+] Task reported complete.")
break
time.sleep(1.0) # Allow application state to settle
4. Deployment, Rate Limits, and Error Handling
1. Frame Buffer Throttling
Sending entire screen frames per token step can exhaust LLM rate limits (especially TPM limits on images). To handle this:
- SSIM Gating: Compute the Structural Similarity Index between consecutive frames. If visual delta is
< 1%, do not re-send an image; tell the LLM that the UI hasn't visually updated yet.
- Sub-region Cropping: If an active window handle is known, crop the capture strictly to the target bounding box rather than the whole 4K desktop.
2. Guardrails & Failsafes
Always ensure native emergency interrupts remain functional:
- Keep
pyautogui.FAILSAFE = True enabled. Slamming the mouse into any screen corner immediately raises pyautogui.FailSafeException and aborts execution.
- Run tasks within isolated sandboxes, dedicated virtual displays (via Xvfb on Linux), or isolated VMs when automating irreversible accounting or ERP entries.
5. Conclusion & Production Blueprint
Selector-free RPA turns brittle, element-bound automations into reliable computer-use pipelines. By treating UI like human operators do—through visual feedback—you can automate legacy software without maintenance overhead or costly commercial runner licenses.
You can implement this architecture using the snippets above, or deploy our complete, production-hardened template:
The production repository includes:
- Full cross-platform window management (X11, macOS Quartz, Win32).
- Local SSIM frame-caching to cut LLM token costs by over 60%.
- Pre-built error recovery routines with step-level visual debugging.
- Complete JSONL schema validators and visual heatmap audit tools.