
nuphus-mcp
by mrpulor-gh
nuphus-mcp
Desktop automation MCP server — computer use for any AI agent. See the screen, control windows/mouse/keyboard, and drive Chrome over the Model Context Protocol (stdio). Desktop & browser automation need no API key; OCR runs locally; vision plugs into your own vision LLM (OpenAI-compatible, BYOK).
nuphus-mcp is a lightweight, cross-platform desktop automation MCP server
that exposes desktop + browser automation as standard MCP tools. It speaks
JSON-RPC 2.0 over stdio — no daemon, no network service, one binary. Claude
Desktop, Cursor, VS Code, Copilot, or any MCP client can connect and
immediately control the screen, windows, keyboard/mouse, and Chrome —
computer use for any AI agent — desktop & browser automation need no API
key; local OCR is built in; vision works with your own vision LLM
(OpenAI-compatible, BYOK).
🇨🇳 Mainland China mirror: this repo is mirrored on Gitee for fast in-China access (Chinese docs served by default there). 中文文档
┌──────────────────┐ stdio JSON-RPC ┌──────────────────────┐
│ Any MCP Client │ ───────────────► │ nuphus-mcp │
│ (Claude/Cursor/ │ ◄─────────────── │ desktop-api crate │──► screen/window/mouse/keyboard
│ Nuphus itself) │ single-line JSON │ nuphus-browser crate│──► Chrome (CDP)
└──────────────────┘ └──────────────────────┘
Features
- 36 MCP tools (15 desktop + 21 browser) — screenshots, window control, mouse/keyboard, Chrome CDP automation, and more — see TOOLS.md / TOOLS.zh-CN.md for the full reference.
- Desktop automation: screen size, screenshot (BMP/base64), window list,
window activate/screenshot/move/resize/info, mouse click/drag/scroll/position,
keyboard input/hotkey, clipboard write/clean — implemented on the
desktop-apicrate (xcap + Win32, no Tauri dependency). - Computer vision pair:
desktop_vision(BYOK — send a screenshot to your own vision model via an OpenAI-compatible API) +desktop_perceive(local OCR with PaddleOCR, models auto-downloaded on first run; optional YOLO icon detection). Used together they give AI agents both semantic understanding and pixel-precise coordinates — the battle-tested vision→perceive flow from the Nuphus desktop app. See TOOLS.md for BYOK env vars, model setup, and the recommended flow. - Browser automation: navigate, snapshot (accessibility tree with
@Nrefs), click, type, exec, scroll, extract, screenshot, evaluate, back/forward, wait_for, cookies get/set/import, upload, tabs, downloads — implemented onnuphus-browser(chromiumoxide CDP). - Zero-cost stdio: no HTTP server, no daemon. The process reads single-line JSON from stdin and writes responses to stdout.
- Safety-first: destructive tools are annotated per the MCP spec; optional strict-confirm mode; path validation for screenshot/upload.
Repository Layout
nuphus-mcp/
├── Cargo.toml # workspace root
├── TOOLS.md / TOOLS.zh-CN.md # 36-tool reference
├── crates/
│ ├── nuphus-mcp/ # MCP Server (this repo's product)
│ ├── nuphus-browser/ # Browser automation core (CDP)
│ └── desktop-api/ # Desktop control core (vendored)
└── ...
Prerequisites
- Rust toolchain (stable) — build from source with Cargo.
- Chrome or Edge — required for browser tools. The server auto-detects an
installed browser; if none is found,
browser_*tools return a clear error. - Windows recommended for full desktop control — see Platform Support below.
Platform Support
| Platform | Browser tools | Desktop tools |
|---|---|---|
| Windows | Full | Full (Win32 API) |
| macOS | Full | Desktop input requires Accessibility permission (System Settings → Privacy & Security → Accessibility) |
| Linux | Available |
Related servers

n8n
Updated todayby n8n-io
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

mcp-server-fetch
OfficialUpdated todayA Model Context Protocol server providing tools to fetch and convert web content for usage by LLMs

@modelcontextprotocol/server-filesystem
OfficialUpdated todayMCP server for filesystem access