Backends
pi-web-agent can use alternate search and fetch backends without changing the public web_explore tool.
Self-hosted options:
- SearXNG for search
- Firecrawl for page fetch/extraction
Hosted options:
- Brave Search for API-backed source discovery
- You.com Search for API-backed source discovery
- Exa for API-backed source discovery
- Tavily for API-backed source discovery
- Google SERP for any hosted vendor that returns Google results (Serper, SerpBase, and similar)
These are all hosted, so they use an API key. Only google-serp also needs a baseUrl, since you choose the vendor.
This keeps the public Pi tool the same: the model still calls web_explore. The backend config only changes what web_explore uses internally.
What this page does not cover
This project does not manage SearXNG or Firecrawl deployments.
Use the upstream docs for:
- installing either service
- Docker Compose files
- reverse proxies
- TLS
- auth setup
- service upgrades
The assumption here is that you already have working services and just want pi-web-agent to connect to them.
Default backend config
Without any backend config, pi-web-agent uses:
{
"backends": {
"search": { "provider": "duckduckgo" },
"fetch": { "provider": "http" },
"headless": { "provider": "local-browser" }
}
}That path does not require SearXNG or Firecrawl.
Keyless default
With no configuration, search uses DuckDuckGo. We send normal browser headers, so it holds up better than a raw scrape. If DuckDuckGo walls the request (common on datacenter IPs), search falls back to Tavily's keyless endpoint (no account, no API key) straight away, without retrying the blocked page. Only the query that failed is sent, and only on failure. If DuckDuckGo answers but simply found nothing, that empty result is returned as is and Tavily is not asked.
Don't want the Tavily fallback? Set PI_WEB_AGENT_DISABLE_KEYLESS_FALLBACK=1 and a blocked DuckDuckGo request will just return an error instead.
Easiest setup path
Open:
/web-agent settingsChoose Backends. From there you can:
- switch search between DuckDuckGo, SearXNG, Brave, You.com, Exa, Tavily, and Google SERP
- edit the search endpoint URL (SearXNG or Google SERP)
- enable a hosted/SearXNG → DuckDuckGo fallback
- switch fetch between plain HTTP and Firecrawl
- edit the Firecrawl base URL
- enable Firecrawl → HTTP fallback
- set the outbound proxy URL
Hosted search and Firecrawl API keys are intentionally not edited in the settings UI. Prefer environment variables for secrets: PI_WEB_AGENT_BRAVE_API_KEY, YDC_API_KEY, EXA_API_KEY, TAVILY_API_KEY, PI_WEB_AGENT_GOOGLE_SERP_API_KEY, and PI_WEB_AGENT_FIRECRAWL_API_KEY.
Config file locations
Global config:
~/.pi/agent/extensions/pi-web-agent/config.jsonProject config:
.pi/extensions/pi-web-agent/config.jsonProject config overrides global config.
SearXNG search
To use SearXNG for search, choose Settings → Backends, set search provider to searxng, and enter the base URL. The equivalent config is:
{
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080"
}
}
}pi-web-agent expects SearXNG JSON search to work at:
/search?q=example&format=jsonRun this after editing config:
/web-agent doctorDoctor checks that the configured SearXNG endpoint responds with JSON that looks like search results.
Supported SearXNG options can stay in config:
{
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080",
"options": {
"categories": ["general", "it"],
"language": "en",
"safesearch": 1
}
}
}
}These map to SearXNG search query params. Config loading filters the options before doctor validates the effective config. Unknown options and some malformed values, such as safesearch: 9 or a categories array containing a non-string, are ignored. Empty categories or a blank language survive loading and produce warnings. backend config: ok therefore does not mean every field in the file was accepted. Whether loading should report every rejected field is still an open question.
Google SERP endpoint
google-serp is a vendor-neutral wrapper around any hosted service that front-ends Google results (Serper, SerpBase, and similar). Nothing in it is tied to one vendor: you set the endpoint and the key, and switching vendors is a base-URL change rather than a config migration.
Set the key:
PI_WEB_AGENT_GOOGLE_SERP_API_KEY=...Then pick Settings → Backends → Search backend → google-serp and enter the endpoint under Search endpoint URL, or write it directly:
{
"backends": {
"search": {
"provider": "google-serp",
"baseUrl": "https://google.serper.dev/search"
}
}
}pi-web-agent POSTs this body to that URL:
{ "q": "example", "num": 10 }with the key in a header (X-API-Key by default) and expects the common organic[] shape back:
{
"organic": [
{ "title": "Example", "link": "https://example.com", "snippet": "..." }
]
}If your vendor spells the header differently, set keyHeader:
{
"backends": {
"search": {
"provider": "google-serp",
"baseUrl": "https://api.example.com/search",
"keyHeader": "Authorization"
}
}
}The key is sent as-is in the header, so if the header needs a scheme (like Authorization: Bearer <key>), include the scheme in the key itself.
Two things worth knowing:
- Vendors in this space often answer HTTP 200 with a status envelope in the body. A 2xx status in the body is treated as success even if it carries a message field. Higher status codes (e.g., 1001 for unauthorized, 1504 for transient) are checked against documented error codes and wording to classify the failure, so a bad key or empty balance shows up as an auth or quota failure instead of "no results".
- SerpApi needs the key as a query parameter and returns
organic_results. That request/response profile is not supported by this adapter.
Run /web-agent doctor after editing config: it sends the same one-result probe and reports search backend: google-serp ok, a warning with the reason, or the missing base URL/key.
Brave Search
To use Brave Search, set:
PI_WEB_AGENT_BRAVE_API_KEY=...Then choose Settings → Backends → Search backend → brave.
Equivalent config:
{
"backends": {
"search": {
"provider": "brave",
"fallback": "duckduckgo"
}
}
}Brave only improves source discovery. web_explore still fetches pages, ranks evidence, handles headless fallback, and writes caveats itself.
You.com Search
To use You.com Search, set:
YDC_API_KEY=...Then choose Settings → Backends → Search backend → youcom.
Equivalent config:
{
"backends": {
"search": {
"provider": "youcom",
"fallback": "duckduckgo"
}
}
}You.com only improves source discovery. web_explore still fetches pages, ranks evidence, handles headless fallback, and writes caveats itself.
Exa
To use Exa, set:
EXA_API_KEY=...Then choose Settings → Backends → Search backend → exa.
Equivalent config:
{
"backends": {
"search": {
"provider": "exa",
"fallback": "duckduckgo"
}
}
}Exa only improves source discovery. web_explore still fetches pages, ranks evidence, handles headless fallback, and writes caveats itself.
Tavily
To use Tavily, set:
TAVILY_API_KEY=...Then choose Settings → Backends → Search backend → tavily.
Equivalent config:
{
"backends": {
"search": {
"provider": "tavily",
"fallback": "duckduckgo"
}
}
}Tavily only improves source discovery. web_explore still fetches pages, ranks evidence, handles headless fallback, and writes caveats itself.
Firecrawl fetch
To use Firecrawl for page reading, choose Settings → Backends, set fetch provider to firecrawl, and enter the base URL. The equivalent config is:
{
"backends": {
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002"
}
}
}pi-web-agent calls Firecrawl's scrape endpoint:
/v1/scrapeIf your Firecrawl instance requires an API key, prefer an environment variable:
PI_WEB_AGENT_FIRECRAWL_API_KEY=...The settings UI does not write API keys. You can still set an API key in config for local-only setups:
{
"backends": {
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002",
"apiKey": "..."
}
}
}Avoid committing project config files that contain secrets.
Saving Backends replaces the selected scope's backend section without a manually stored Firecrawl apiKey or proxy password, including when you save an unrelated field. Presentation saves preserve the backend section. Use environment variables if you want those secrets to survive backend settings saves.
Supported Firecrawl options can stay in config:
{
"backends": {
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002",
"options": {
"formats": ["markdown"],
"onlyMainContent": true
}
}
}
}These are sent in the Firecrawl scrape request body. The supported set is intentionally small for now.
pi-web-agent checks the page URL before sending it to Firecrawl and refuses private addresses. What a Firecrawl server fetches on its own side, such as redirects it follows, is outside that check.
Search fanout
Search fanout queries several configured search providers at once, dedupes the merged results by URL, and reranks so URLs multiple providers agree on rank higher. Then the normal research loop continues.
Fanout has three modes:
off: Default. Search uses a single provider (configured via Settings → Backends → Search backend).on: Every search queries all configured providers and merges the results.auto: Runs the primary provider first. If its results look thin (too few, or all from one host), fans out to configured providers for better coverage.
When fanout runs, the provider set defaults to the providers you have actually configured: DuckDuckGo (always, keyless), SearXNG if you gave it a baseUrl, and each hosted provider only when its API key is set. Providers you have not set up are not offered or queried.
To select which providers fan out, set the fanout mode from Settings → Backends, then toggle individual providers on or off. Only usable providers appear in that list. You can also edit backends.search.fanout.providers directly in the config file.
The selected primary provider always participates, even if its toggle says excluded or you omit it from fanout.providers. A configured DuckDuckGo fallback is also included under fanout, even when it is absent from that list. Disabling fallback stops this automatic DuckDuckGo addition; excluding the primary requires selecting a different primary. An omitted or empty provider list uses all usable providers.
Each provider gets a short timeout during fanout, so one slow or unreachable backend (for example a self-hosted SearXNG that is down) is skipped instead of stalling the whole research pass.
searxng and google-serp need endpoints. For the selected endpoint provider, backends.search.baseUrl takes precedence over its backends.search.baseUrls.<provider> entry. Other providers use their own entries. Legacy configs with DuckDuckGo or a hosted provider selected can still use baseUrl for SearXNG; Google SERP does not reuse that URL. A fanout set that mixes both endpoint providers looks like this:
{
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080",
"baseUrls": { "google-serp": "https://google.serper.dev/search" },
"fanout": { "mode": "on", "providers": ["duckduckgo", "searxng", "google-serp"] }
}
}
}A provider with no endpoint of its own stays out of the set instead of being pointed at the other one's URL, so your Google key never travels to SearXNG (or the other way round). /web-agent doctor names the key each provider is missing.
The equivalent config is:
{
"backends": {
"search": {
"provider": "brave",
"fanout": { "mode": "auto", "providers": ["duckduckgo", "brave", "exa"] }
}
}
}In preview and verbose modes, fanout visibility shows which providers were queried, for example:
web_search ×2 (fanout: duckduckgo, brave, exa)Trade-off: fanout costs extra latency and API calls, which is why it is off by default and auto exists. Use on when you want maximum source diversity at the cost of longer research times. Use auto when you want a safety net without paying the latency cost most of the time.
Explicit fallback
Fallback is opt-in. pi-web-agent does not silently leave a self-hosted backend unless you configure it. You can turn fallback on from Settings → Backends. The equivalent config is:
{
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080",
"fallback": "duckduckgo"
},
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002",
"fallback": "http"
}
}
}When fallback happens, output indicates which backend failed and which fallback was used. This keeps self-hosted privacy expectations explicit: if you do not configure fallback, SearXNG, Brave, You.com, Exa, Tavily, Google SERP, and Firecrawl failures stay visible instead of silently switching to external/default backends.
Fallback also looks at why a backend failed before moving on:
- A rate limit falls back and skips that backend for a while: as long as the provider asked, capped at 15 minutes, or a minute if it didn't say.
- An exhausted quota, a rejected API key, or a missing key or base URL falls back and skips that backend until your settings change.
- A timeout, server error, or dropped connection gets one retry first.
- A blocked or garbled response falls back without a retry.
- An empty result is a real answer, so it does not fall back. Neither does a bad request (like an empty query).
- A private-address refusal or an invalid shared proxy setting never falls back to another backend.
Verbose output lists every backend that was tried, retried, or skipped and why. When some backends failed but another one answered, the answer notes that results may be incomplete. If every backend is being skipped, see "Search says no backend is available" in the troubleshooting guide.
Proxy
Route all outbound web_explore traffic through an HTTP proxy:
{
"backends": {
"proxy": {
"url": "http://127.0.0.1:7890",
"username": "user",
"password": "pass"
}
}
}urlis required and must be anhttp://orhttps://proxy URL. It must not embed credentials (http://user:pass@host:port); pi-web-agent strips anyuser:pass@from the URL and never sends it. Put credentials inusername/passwordor the environment variables below instead.usernameandpasswordare optional. When present, credentials are sent to the proxy as aProxy-Authorizationheader (never to the target site).- Credentials can instead come from the environment variables
PI_WEB_AGENT_PROXY_USERNAMEandPI_WEB_AGENT_PROXY_PASSWORD. Config values win when both are set. If you set credentials in the URL,/web-agent doctorand config validation tell you to move them to these variables.
When a proxy is configured it applies to every outbound request: search backends, plain HTTP fetches, Firecrawl, the GitHub and YouTube readers, PDF downloads, the headless browser, and /web-agent doctor health checks. HTTPS targets are tunneled with CONNECT, so the proxy never sees the request contents.
The proxy URL is editable from Settings → Backends. The settings UI does not edit proxy credentials; like other secrets, keep credentials in environment variables rather than committed config files.
A proxy that is unreachable makes requests fail with a clear proxy error instead of silently bypassing the proxy.
Page fetches, the readers, and the headless browser don't talk to your proxy directly. They go through a small local guard proxy that looks up each address, refuses private ones, and then asks your proxy to connect to the IP it checked. So for those requests your proxy sees an IP address rather than the hostname; search backends and Firecrawl still reach it by hostname. If your proxy only accepts hostnames, see trustProxyDns below and "Fetches fail with an upstream proxy refused error" in the troubleshooting guide.
Private addresses, allow list, and proxy trust
web_explore refuses to connect to private, loopback, and link-local addresses (your LAN, localhost, cloud metadata like 169.254.169.254) when the link came from the model or from a page it read. It checks every address a name resolves to, including redirects and everything a headless page loads, and connects to the address it checked so a DNS change can't move the request. Addresses you configured yourself are not affected: search backends, SearXNG, Firecrawl, and the proxy.
Two settings adjust this, both under Settings → Backends:
{
"backends": {
"network": {
"allowRanges": ["10.0.0.0/24"],
"trustProxyDns": false
}
}
}allowRangeslists CIDR rangesweb_exploremay reach anyway, for example an internal docs site. Invalid entries and ranges that allow everything (0.0.0.0/0,::/0) are rejected. If you run a proxy app in fake-IP mode, every site resolves inside198.18.0.0/15; add that range or every fetch will be refused.trustProxyDnshands hostnames to your upstream proxy instead of checked IPs. Turn it on for proxies that only accept hostnames, or networks where names only resolve inside the proxy. It trusts your proxy to keep requests away from private addresses, so it is off by default, and localhost, private IPs written into a link, and names your own machine resolves to a private address are still refused. It has no effect withoutbackends.proxy.
A project config's allow list replaces the global one rather than adding to it. /web-agent doctor shows both settings.
The protection covers connections that go through the guard proxy. It is not a full network sandbox for the browser; if you need that, restrict outbound traffic at the OS or container level.
Full self-hosted example
{
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080",
"fallback": "duckduckgo",
"options": {
"categories": ["general"],
"language": "en",
"safesearch": 1
}
},
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002",
"fallback": "http",
"options": {
"formats": ["markdown"],
"onlyMainContent": true
}
},
"headless": {
"provider": "local-browser"
}
}
}You can combine this with presentation settings in the same file. The settings UI preserves both sections when saving:
{
"presentation": {
"defaultMode": "preview"
},
"backends": {
"search": {
"provider": "searxng",
"baseUrl": "http://localhost:8080",
"fallback": "duckduckgo"
},
"fetch": {
"provider": "firecrawl",
"baseUrl": "http://localhost:3002",
"fallback": "http"
},
"headless": {
"provider": "local-browser"
}
}
}Verify the setup
Show the effective config:
/web-agent showRun diagnostics:
/web-agent doctorExpected healthy output includes lines like:
search: searxng (http://localhost:8080) fallback duckduckgo
fetch: firecrawl (http://localhost:3002) fallback http
backend config: ok
search backend: searxng ok
search fallback: duckduckgo
fetch backend: firecrawl ok
fetch fallback: http
headless backend: local-browser (managed Chromium fallback configured)Then try a normal research prompt:
Find current docs for configuring Vitest coverage with the v8 provider.The model should still use web_explore; it should not need separate SearXNG, Brave, You.com, Exa, Tavily, Google SERP, or Firecrawl tool calls. If your prompt includes an HTTP/HTTPS URL, web_explore reads that URL before spending search passes.
Troubleshooting
search provider searxng requires backends.search.baseUrl
You set provider to searxng but did not include baseUrl.
fetch provider firecrawl requires backends.fetch.baseUrl
You set provider to firecrawl but did not include baseUrl.
search backend: searxng warning
Check that:
- SearXNG is running
- the configured URL is reachable from the Pi process
- JSON output works with
format=json
fetch backend: firecrawl warning
Check that:
- Firecrawl is running
/v1/scrapeis available- the API key is set if your instance requires auth
- the Pi process can reach the configured URL
Self-hosted privacy expectations
pi-web-agent does not silently fall back from SearXNG, Brave, You.com, Exa, Tavily, or Google SERP to DuckDuckGo, or from Firecrawl to plain HTTP, when you choose those providers. Fallback only happens when fallback is configured because some users choose specific backends to control where requests go.
Clear inherited backend settings
Project settings inherit the global config. Clearing an optional value in Settings → Backends saves a backends.cleared marker in the project file so the value stays cleared after reload.
For example, this project config clears the inherited SearXNG endpoint, whether the global config uses baseUrl or baseUrls.searxng:
{
"backends": {
"cleared": ["search.endpoints.searxng"]
}
}When the provider stays the same, omitting a setting inherits it. Changing the search provider drops the lower layer's legacy baseUrl, fallback, options, and fanout configuration unless this layer supplies them; the per-provider baseUrls map and keyHeader still inherit. Changing the fetch provider drops its lower-layer URL, key, fallback, and options. Per-provider search endpoint maps merge across layers, so setting one entry leaves the other inherited.
A clear marker removes a value before this layer's values are applied, so an explicit value in the same layer or a higher layer can restore it. Entering a new URL in settings replaces the cleared endpoint. Removing the marker, or resetting project config, restores inheritance.
Settings saves use search.endpoints.searxng or search.endpoints.google-serp to keep that provider cleared if the global config changes URL formats. Each marker leaves the other provider's endpoint inherited. The field paths search.baseUrl and search.baseUrls.<provider> clear only the named field.
Clear markers support optional backend URLs, per-provider endpoints, headers, fallbacks, options, fanout configuration, proxy credentials and network settings. Required provider fields cannot be cleared. Unknown paths are ignored. Clearing the selected provider's required endpoint leaves it unconfigured; choose another provider or enter a new endpoint before using it.