Make the three Browser Use paths consistent between the diagram and quickstarts: fully hosted cloud, CLI, and Python library. Combine the open source Browser Use agent and its Python library into one diagram block connected directly to local and cloud browsers. Rename the demo heading to “Navigate the web like a human does.” and mention choosing a date and time. Keep the inline animation, remove the separate Johannes link, expand the hosted-service description, and remove the no-card text and Odysseys claim. Validation: pre-commit, whitespace checks, Python snippet parsing, numbered section order and anchors, SVG parsing, and single-block structure in both diagram themes passed. Inspected both diagram themes and the actual GitHub README on desktop and at 390px mobile width. All three quickstart anchors resolve, the inline GIF remains present, and there is no page overflow.
What can Browser Use do?
Browser Use gives AI agents a browser. Describe a task, and the agent navigates websites, fills forms, and gets the work done.
Navigate the web like a human does.
Find an available slot, pick a date and time, handle the CAPTCHA, and book a driving test.
Explore more demos and prompts ↗
Which Browser Use do I need?
- Path 1: Fully Hosted Cloud: Send a task through the API; we run the agent and browser.
- Path 2: CLI: Give Claude Code, Codex, OpenCode, Pi, or another agent browser access.
- Path 3: Python Library: Run the open source Browser Use agent locally from your own code.
Quickstart
Path 1: Fully Hosted Cloud
Send a task to our hosted agent. We run the agent, stealth browser, and scalable infrastructure, and handle profiles, recordings, and data policies so you can scale your browser automation.
New Google, GitHub, or Microsoft signups get $15 cloud credit.
Path 2: CLI
If you want to use Browser Use in your agent (Claude Code, Codex, OpenCode, Pi, Cursor, Hermes, OpenClaw, etc.), paste this prompt, and it sets everything up itself:
Install or upgrade browser-use to the latest stable version with uv using Python 3.12, run `browser-use skill install` to register the skill, and connect it to my browser. If setup or connection fails, follow https://github.com/browser-use/browser-harness/blob/main/install.md.
Then tell your agent what you want done.
Path 3: Python Library
Run the Browser Use agent locally from Python, with your choice of model and a local or cloud browser:
1. Install Browser Use (Python >= 3.11):
uv add browser-use
# or: pip install browser-use
2. Add your OpenAI API key to .env:
# .env
OPENAI_API_KEY=your-key
3. Run your first agent:
import asyncio
from browser_use import Agent, ChatOpenAI
from dotenv import load_dotenv
load_dotenv()
async def main():
agent = Agent(
task="Find the number of stars of the browser-use repo",
llm=ChatOpenAI(model='gpt-5.6-luna', reasoning_effort='xhigh'),
)
history = await agent.run()
if __name__ == "__main__":
asyncio.run(main())
Benchmark
This very hard benchmark targets the hardest browser tasks. On easier tasks, even smaller models can achieve very high success rates. Results shown are from a 60-task subset of BU Bench V2.
Integrations, hosting, custom tools, MCP, and more on our Docs ↗
FAQ
Should I use the CLI vs. the Python library?
Use the CLI if you already have an agent (Claude Code, Codex, OpenCode, Pi, Cursor, Hermes, OpenClaw, etc.) that you want to complete browser tasks for you. The agent installs the skill once (see CLI quickstart) and can then control the browser. Examples:
- "Upload this video to YouTube"
- "Compare these three laptops and give me a table with prices"
- "Fill in this job application with my resume"
Use the Python library when you are building software that automates the web. Examples:
- Run many tasks on a schedule or in parallel (scraping, monitoring, QA)
- Embed a browser agent into your own product
- Custom tools, custom system prompts, structured output, fine-grained browser control
Rule of thumb: one-off tasks through an agent → CLI. Repeatable automation in code → Python library.
What's the best model to use?
We optimized ChatBrowserUse() specifically for browser automation tasks. On avg it completes tasks 3-5x faster than other models with SOTA accuracy.
For pricing and other LLM providers, see our supported models documentation.
Can I use Claude / GPT / Gemini through ChatBrowserUse?
Yes. ChatBrowserUse accepts provider-prefixed model ids, so a single BROWSER_USE_API_KEY reaches all of them — no separate OpenAI/Anthropic/Google keys required:
from browser_use import Agent, ChatBrowserUse
llm = ChatBrowserUse(model='anthropic/claude-sonnet-4-6') # or 'google/gemini-3-pro'
agent = Agent(task='...', llm=llm)
For the best speed and cost we still recommend the default bu-* models.
Should I use the Browser Use system prompt with the open-source preview model?
Yes. If you use ChatBrowserUse(model='browser-use/bu-30b-a3b-preview') with a normal Agent(...), Browser Use still sends its default agent system prompt for you.
You do not need to add a separate custom "Browser Use system message" just because you switched to the open-source preview model. Only use extend_system_message or override_system_message when you intentionally want to customize the default behavior for your task.
If you want the best default speed/accuracy, we still recommend the newer hosted bu-* models. If you want the open-source preview model, the setup stays the same apart from the model= value.
Can I use custom tools with the agent?
Yes! You can add custom tools to extend the agent's capabilities:
from browser_use import Tools
tools = Tools()
@tools.action(description='Description of what this tool does.')
def custom_tool(param: str) -> str:
return f"Result: {param}"
agent = Agent(
task="Your task",
llm=llm,
browser=browser,
tools=tools,
)
Can I use this for free?
Yes! Browser-Use is open source and free to use. You only need to choose an LLM provider (like OpenAI, Google, ChatBrowserUse, or run local models with Ollama).
Terms of Service
This open-source library is licensed under the MIT License. For Browser Use services & data policy, see our Terms of Service and Privacy Policy.
How do I handle authentication?
Check out our authentication examples:
- Using real browser profiles - Reuse your existing Chrome profile with saved logins
- If you want to use temporary accounts with inbox, choose AgentMail
- To sync your auth profile with a remote browser, install
profile-usefor your platform from the official releases, then follow the profile sync guide.
These examples show how to maintain sessions and handle authentication seamlessly.
How do I solve CAPTCHAs?
For CAPTCHA handling, you need better browser fingerprinting and proxies. Use Browser Use Cloud which provides stealth browsers designed to avoid detection and CAPTCHA challenges.
How do I go into production?
Chrome can consume a lot of memory, and running many agents in parallel can be tricky to manage.
For production use cases, use our Browser Use Cloud API which handles:
- Scalable browser infrastructure
- Memory management
- Proxy rotation
- Stealth browser fingerprinting
- High-performance parallel execution
Citation
If you use Browser Use in your research or project, please cite:
@software{browser_use2024,
author = {Müller, Magnus and Žunič, Gregor},
title = {Browser Use: Enable AI to control your browser},
year = {2024},
publisher = {GitHub},
url = {https://github.com/browser-use/browser-use}
}