Agent commands
extract
Section titled “extract”webx extract <url> — page → schema’d JSON. Three modes, cheapest first:
| Mode | Flag | LLM needed? |
|---|---|---|
| Deterministic CSS | --css '{"title":"h1","links":"a[]@href"}' |
No — selector extraction |
| schema.org entities | --type product|article|job|event|faq |
No — reads JSON-LD |
| LLM schema | default path | Yes (WEBX_LLM_*) |
webx extract <url> --css '{"title":"h1","price":".price","imgs":"img[]@src"}'webx extract <url> --type productwebx extract <url> --render # JS pagesPrefer --css when you know the shape — it’s deterministic, free, and reproducible.
webx mcp — run webx as an MCP server over stdio. On a running serve/webxd, the same 17 tools are exposed over streamable HTTP at POST /mcp.
Tools: search · scrape · ask · query · map · doctor · crawl · extract · diff · research · verify · watch · similar · wayback · llms · answer · batch
Client setup: Agents, MCP & skills.
webx skill — print a SKILL.md that teaches agent harnesses (Claude Code, Devin, Cursor, Windsurf) when to use webx and how to call it:
webx skill > .devin/skills/webx/SKILL.mdbotkey
Section titled “botkey”webx botkey — generate a Web Bot Auth ed25519 key + the JWKS to self-host. Set WEBX_BOT_KEY + WEBX_BOT_DIRECTORY and every outbound request is cryptographically signed (RFC 9421, draft-meunier-webbotauth) — a verifiable identity for your crawler instead of fingerprint-guessing. The anti-stealth: sites that adopt the spec can recognize and allow your bot.