Integration examples and documentation for the CapSolver Python ecosystem.
This repo is a documentation and examples hub. The CapSolver packages are installed from PyPI; local checkouts of the package repos are only useful as reference source.
This hub connects three standalone packages:
| Repo | Package | Description |
|---|---|---|
| capsolver-core | capsolver-core |
Core SDK — detect captchas, solve via API, autofill tokens |
| capsolver-agent | capsolver-agent |
Agent tools — framework-agnostic schemas + LangChain BaseTool |
| capsolver-mcp | capsolver-mcp |
MCP Server — expose captcha tools via Model Context Protocol |
# Install the packages you need
pip install capsolver-core
pip install capsolver-agent
pip install capsolver-mcp
# Browser-aware examples use optional extras
pip install "capsolver-core[playwright]"
pip install "capsolver-agent[browser,langchain]"
pip install "capsolver-mcp[browser]"
playwright install chromium
# Set your API key
# bash / zsh
export CAPSOLVER_API_KEY="your-capsolver-api-key"
# PowerShell
$env:CAPSOLVER_API_KEY = "your-capsolver-api-key"
# cmd
set CAPSOLVER_API_KEY=your-capsolver-api-keyFor runnable examples, copy the reference environment file and fill in local values:
cp .env.example .envPowerShell:
Copy-Item .env.example .envLeave OPENAI_BASE_URL unset to use the official OpenAI endpoint. Set it only
when testing an OpenAI-compatible gateway such as OpenRouter or DeepSeek.
Runnable demos showing how to integrate CapSolver into popular AI frameworks:
| File | Framework | Packages | Default behavior |
|---|---|---|---|
openai_function_calling.py |
OpenAI API | capsolver-agent |
LLM lists supported captcha types; no solve by default |
openai_agents.py |
OpenAI Agents SDK | capsolver-agent |
LLM lists supported captcha types; no solve by default |
langchain_agent.py |
LangChain | capsolver-agent |
LLM lists supported captcha types; no solve by default |
browser_use_agent.py |
Browser Use | capsolver-agent, capsolver-core |
LLM drives an RPA captcha workflow on the active Browser Use page |
playwright_sdk.py |
Playwright | capsolver-core |
No LLM; direct SDK detect/solve/fill workflow |
See examples/README.md for per-example setup and running instructions.
# Tool/schema smoke tests without LLM calls
python examples/openai_function_calling.py --list-tools
python examples/openai_agents.py --list-tools
python examples/langchain_agent.py --list-tools
# Read-only LLM examples
python examples/openai_function_calling.py
python examples/openai_agents.py
python examples/langchain_agent.py
# Browser Use: trigger captcha, solve on the active Browser Use page, submit
python examples/browser_use_agent.py --mode agent --url "https://www.google.com/recaptcha/api2/demo" --no-headless --captcha-timeout 180 --step-timeout 240
# Playwright direct SDK: trigger, solve, fill, submit
python examples/playwright_sdk.py --url "https://www.google.com/recaptcha/api2/demo" --trigger-captcha --submit-after-fill --no-headless --captcha-timeout 180MCP clients such as Claude Code or Codex CLI do not need a Python demo.
Configure capsolver-mcp, then ask the client to use the exposed tools:
codex mcp add --env CAPSOLVER_API_KEY=CAP-XXXXXX capsolver -- uvx capsolver-mcp
codex "Use the capsolver MCP tools to list supported captcha types. Do not solve a captcha or check account balance."- LLM-powered examples use the LLM only for workflow decisions and tool selection. Captcha solving is delegated to CapSolver tools.
capsolver-agentandcapsolver-mcpURL tools such assolve_on_pageopen a browser owned by that tool call. They do not fill an arbitrary existing RPA browser page.examples/browser_use_agent.pyincludes a Browser Use specificauto_solve_current_page_captchasaction that adapts the activeBrowserSessiontocapsolver-core's page-driver interface. Use that pattern when the captcha must be filled on the same page the RPA agent is controlling.- Many RPA pages require a normal user action before a captcha appears. The Browser Use and Playwright examples both support clicking common human-verification entry points before solving.
- MCP clients such as Claude Code, Codex CLI, Cursor, Windsurf, and Cline parse the MCP tool schemas themselves. They usually need configuration and prompt examples, not a separate Python script.
- docs/integrations.md — Integration methodology: architecture, choosing the right layer, mapping schemas, framework-specific patterns.
- docs/agent-integration.md — Framework integration guides: OpenAI, OpenAI Agents SDK, LangChain, LlamaIndex, CrewAI, Google ADK, Mistral, Vercel AI SDK, and custom frameworks.
- docs/mcp-integration.md — MCP client setup: Claude Desktop, Claude Code, Codex CLI, Cursor, Windsurf, Cline, and remote HTTP mode.
Each package ships a CLI for development and debugging:
# SDK diagnostics
capsolver info # version, Python, optional deps
capsolver list-types # supported captcha types
capsolver balance # account balance
# Agent tool inspection
capsolver-agent list # all tools with descriptions
capsolver-agent schema solve_captcha --format openai
# MCP server
capsolver-mcp # stdio (for Claude Desktop, Cline, etc.)
capsolver-mcp --transport sse # SSE for remote accessgit clone https://github.com/pandapro-project/capsolver-ai-hub.git
cd capsolver-ai-hub
pip install -r requirements.txt
pip install -r requirements-dev.txtEach package has its own test suite in its package repo. This hub focuses on example scripts and documentation that consume the PyPI packages.
MIT