Forked from
Forked from
Keywords: lm studio plugin, web search ai, local web research, fact checking ai, source verification, no api key, offline capable, private browsing ai
A research-grade web search plugin that goes beyond snippets โ it reads pages, verifies claims, detects contradictions, finds primary sources, and enforces cross-source fact verification before asserting anything as true.
Standard search gives you ten links and short snippets. You get:
This plugin fixes all of that.
Load the built plugin in LM Studio.
| Field | Default | Description |
|---|---|---|
| Max Search Results | 8 | Results retrieved per query |
| Max Pages to Read | 3 | Pages actually fetched and read per search |
| Page Fetch Timeout | 8000ms | Per-page timeout before giving up |
| Search Interval | 4000ms | Minimum interval between search requests (all enabled providers). Higher = fewer CAPTCHA blocks, slower. If blocked often, raise to 5000โ6000 ms or use SearXNG. |
| SearXNG | OFF | Self-hosted first tier, opt-in (compact toggle; details on hover). Most users do not host SearXNG โ leave off unless you have a URL. |
| DuckDuckGo | ON | Primary DDG scraper (compact). |
| Bing | ON | Fallback tier (compact). |
OFF | Google scraper, stricter bot protection (compact). | |
| Ecosia | OFF | Bing-powered, privacy (compact). |
| Startpage | ON | Google proxy, privacy (compact). |
| [Dubious] Yandex | OFF | Russian, heavy bot protection (compact). |
| [Dubious] Yahoo | OFF | Bing-powered, ad-heavy (compact). |
| [Dubious] Baidu | OFF |
Before running any search, the plugin calls a clarify step. It detects ambiguity signals in the user's question:
If ambiguity is detected, the LLM asks the user focused questions before searching. This avoids wasted searches and produces a much more targeted answer.
After retrieving search results, the plugin reranks them using embeddings before fetching any pages. This means the pages that actually get read are the most semantically relevant to the query โ not just the top SEO results.
How it works:
nomic-embed-text (requires the model loaded in LM Studio)If the embedding call fails (model not loaded, wrong URL), it falls back silently to the original search engine ranking.
search, search_recent, and search_news each report independent_publishers_read โ the number of distinct root domains among the pages successfully fetched. This feeds a dynamic instruction injected into every result:
| Publishers read | Instruction to the LLM |
|---|---|
| 0 | Do not assert any facts โ re-search or inform the user |
| 1 | Hard UNVERIFIED warning โ call fact_check or label every claim as unverified before presenting |
| 2+ | Report publisher count โ flag any claim supported by only one of them as UNVERIFIED |
This prevents the LLM from repeating a claim as fact just because one website said it.
The system prompt enforces five non-negotiable rules on top of the publisher diversity signal:
Results come from SearXNG (if configured) โ DuckDuckGo โ Startpage (Google proxy, default ON) โ Google (opt-in) โ Ecosia (Bing-powered, opt-in) โ Bing โ [Dubious, opt-in] Yandex / Yahoo / Baidu. No API keys required. All tiers are individually toggleable in settings and share the configurable interval + jitter + randomized User-Agent.
clarify โ Ambiguity check (called automatically first)Always called before any search. Returns either:
STATUS: READY โ question is specific enough, search proceedsSTATUS: CLARIFY โ question is ambiguous, LLM asks user before searchingYou do not need to call this manually. The system prompt enforces it as a mandatory first step.
search โ Core search with page readingThe main tool. Unlike basic search, it fetches and reads the actual page content โ not just snippets.
Returns:
independent_publishers_read โ count of distinct root domains among successful page readsfetch_and_read โ Read a specific URLFetch any URL and return the full readable text content.
Use when:
deep_search โ Multi-angle researchRuns 3โ5 separate searches from different perspectives on the same topic, reads pages for each, and returns everything together. Defeats single-search bias.
Default angles: overview facts, latest research, criticism/limitations, expert consensus.
You can specify your own angles, e.g.:
fact_check โ Verify a specific claimCross-checks a claim across four search angles: direct confirmation, debunking searches, evidence searches, and expert opinion. Returns raw evidence from all angles for the LLM to assess.
Verdict categories: supported, disputed, unsupported, nuanced, uncertain.
verify_statistic โ Verify a number or percentageStatistics are frequently outdated, misquoted, out of context, or fabricated. This tool searches for the stat, its primary source, fact-check results, and updated data.
Example: verify_statistic("90% of startups fail in year one", "venture-backed US tech startups")
find_primary_source โ Trace a claim to its originSecondary sources often distort original findings. This tool searches for the original study, report, official document, or statement where a claim first appeared.
Prioritises: peer-reviewed journals, government reports, official organisation publications over secondary citations.
search_recent โ Time-filtered searchOnly returns results from the specified time window. Prevents stale results from dominating on fast-moving topics.
Windows: day (last 24h), week, month, year.
Returns independent_publishers_read and a dynamic publisher diversity instruction alongside the results.
compare_sources โ Surface agreements and conflictsFetches multiple sources on the same topic and returns them side by side for the LLM to compare framing, spot conflicts, and identify unique claims.
Provide specific URLs to compare, or let it search and pick sources with varied domains automatically.
Returns structured analysis of:
find_expert_views โ Expert consensus and dissentSearches specifically for academic research, official positions, expert interviews, and scientific consensus โ not what random blogs claim experts say.
Covers four angles: expert consensus, peer-reviewed research, official institutional positions, and active scientific debate.
search_academic โ Academic papers onlySearches arXiv, PubMed, and Semantic Scholar for peer-reviewed papers and research publications.
Sources: arxiv, pubmed, semantic_scholar, all.
Fetches paper pages to extract abstracts, methodology, and findings. The LLM is instructed to distinguish preprints from peer-reviewed work, note sample sizes, and not overstate findings.
search_news โ News-specific searchNews-filtered search that actively ranks established journalism above blogs, product pages, and content farms. Runs two queries โ one general, one targeting major news outlets โ then ranks high-credibility results first.
Windows: day, week, month, any.
Unlike search_recent (which filters by date), this filters by source type โ it's about journalistic sourcing, not just recency. Best for: breaking news, corporate announcements, policy changes, anything where "who is reporting it" matters.
Returns independent_publishers_read and a dynamic publisher diversity instruction alongside the results.
research_topic โ Full multi-step research briefRuns multiple searches from different angles, reads key pages, and instructs the LLM to produce a structured research brief: overview, established facts, contested areas, expert consensus, open questions, key sources, and confidence assessment.
Depths:
overview โ 3 angles, 2 pages eachdetailed โ 5 angles, 2 pages each (default)comprehensive โ 7 angles, 3 pages eachcheck_source โ Source credibility assessmentAssesses a URL or domain and returns its credibility type, known signals, reputation search results, and red flags to watch for.
Domain types: government, academic institution, academic/research platform, established news outlet, encyclopedia, user-generated content, unknown.
Credibility levels: high, medium, low, unknown.
Red flags checked:
Every search result and fetched page gets a credibility assessment based on domain signals:
| Domain Type | Credibility | Examples |
|---|---|---|
| Government | HIGH | .gov, .mil, WHO, CDC |
| Academic institution | HIGH | .edu, .ac.uk, universities |
| Academic platforms | HIGH | arXiv, PubMed, Semantic Scholar |
| Established news | HIGH | Reuters, AP, BBC, Nature, NYT |
| Wikipedia | MEDIUM | Good overview, verify citations |
| User-generated / blogs | LOW | Blogspot, WordPress, Reddit, Quora |
| Unknown | UNKNOWN | Check About page and author credentials |
The plugin's system prompt instructs the LLM to:
Simple fact:
"What is the Dunning-Kruger effect?" โ
clarify(READY) โsearch
Ambiguous query:
"Tell me about python" โ
clarify(CLARIFY) โ asks: "Do you mean the programming language or the snake?" โsearch
Verify a claim:
"Is it true that we only use 10% of our brains?" โ
clarify(READY) โfact_check
Verify a statistic:
"Someone told me 50,000 species go extinct every year. Is that right?" โ
clarify(READY) โverify_statistic
Recent developments:
"What's happened with GPT-5 in the last week?" โ
clarify(READY) โsearch_recent(window: "week")
Compare perspectives:
"What do different sources say about seed oils and health?" โ
clarify(READY) โcompare_sourcesordeep_search
Scientific consensus:
"What does the research actually say about intermittent fasting?" โ
clarify(READY) โfind_expert_views+search_academic
Deep research:
"Give me a thorough research brief on quantum error correction" โ
clarify(READY) โresearch_topic(depth: "comprehensive")
Read a specific article:
"Can you read this paper and summarise the key findings? [url]" โ
clarify(READY) โfetch_and_read
Check if a source is reliable:
"Is naturalhealth365.com a reliable source?" โ
clarify(READY) โcheck_source
AI capability claim:
"I read that [model X] achieves 95% accuracy on [benchmark]. Is that right?" โ
clarify(READY) โfact_checkโ vendor blogs and press releases are not accepted as evidence; requires independent academic or journalistic confirmation
Other altra plugins that include web search as a secondary capability will automatically defer to this plugin when it is installed alongside them. Their duplicate search/fetch tools are omitted at startup so this plugin's richer versions take over:
| Plugin | Tools deferred to web-search |
|---|---|
altra/research | search_sources, read_source |
altra/ideas | research |
altra/high-perf-tools | fetch_url, search_web |
Installing this plugin is the recommended way to get the best search quality across all plugins at once. Each affected plugin also shows a tip on first message when this plugin is not installed.
The ddgSearch() function had multiple flaws that caused search results to silently return empty after one or two successful calls per session:
No TLS fingerprint spoofing โ DuckDuckGo performs TLS fingerprinting to detect bots. The native call lacks a realistic TLS fingerprint, so DDG identifies it as a bot and serves a CAPTCHA challenge page ("Select all squares containing a duck") instead of search results. This is the primary cause of persistent empty results.
got-scraping for anti-bot evasion (new dependency):
All search backend requests and page fetches now use got-scraping with headerGeneratorOptions (Chrome 120+, desktop, Linux, en-US). This generates realistic browser headers and TLS fingerprints that bypass DDG's bot detection. Native fetch() is used as a fallback only.
CAPTCHA detection (isDuckDuckGoCaptcha):
Detects DDG's bot challenge pages by checking for anomaly-modal, cc=botnet, challenge-form, and related markers. When detected, throws a clear error instead of silently returning empty results.
Rate limiter (makeRateLimiter + searchRateLimiter):
Enforces a 2000ms minimum interval between search backend requests, preventing DuckDuckGo from blocking or rate-limiting the plugin.
Retry logic (DDG_MAX_RETRIES = 3, DDG_RETRY_WAIT_MS = 2500):
Retries DuckDuckGo up to 3 times with 2.5s waits when results are empty. CAPTCHA errors and network failures fall through immediately to Bing without retrying.
Diagnostic reporting (SearchDiagnostic interface):
Every search call now returns a search_diagnostics array alongside results. Each entry tracks:
backend โ which search engine was tried (searxng, ddg, bing)ok โ whether it returned resultsAll 13 tool implementations updated:
Each tool now includes search_diagnostics in its JSON response, so the model and user can see exactly which backend succeeded or failed and why.
When searches fail, instead of silently empty results, the response includes diagnostics:
This tells you exactly which backends were tried, how many attempts were made, and what went wrong โ so you and the model can diagnose and adapt.
search_recent over-triggering (model defaulted to time-filtered search)Previously the model invoked search_recent/search_news too often even for evergreen queries. Three guardrails were added:
Category split (config.ts:33-68, toolsProvider.ts:92-100,1067-1085): Search providers are now grouped into Regular (SearXNG, DDG, Bing, Google, Ecosia, Startpage) and [Dubious] (Yandex, Baidu, Yahoo). Dubious engines are those with heavier bias/censorship, ad-heavy or less reliable indexes, and aggressive anti-scraping. Each has an individual boolean toggle in the GUI (enableYandex / enableYahoo / enableBaidu โ all OFF by default, opt-in only) with subtitle warnings. Regular toggles are enableEcosia (OFF by default) and enableStartpage (ON by default) in addition to the existing enableSearxng/enableDDG/enableBing/enableGoogle. The unified chain is now ; disabled tiers emit diagnostics. Documents the new toggles; documents the full order.
All new scrapers reuse the hardened fetchSearchHtml() (TLS spoof, jitter interval, randomized UA from userAgents, proper Referer) and report backend: "ecosia"|"startpage"|"yandex"|"yahoo"|"baidu" diagnostics with CAPTCHA/empty handling consistent with DDG/Google/Bing.
promptPreprocessor.ts:74-81, toolsProvider.ts:39-62,173-195,2170-2245Squashed several subtle loops that could still trap the model even after the earlier re-search fix:
Thanks to Altra for making something cool, and giving it to the world to toy around!
Additional "Thank you!" to brius and nub235 for field expertise; The engineering solutions in your code served as a reference implementation for fixes here. Your use of got-scraping for anti-bot evasion, structured error reporting, and rate limiting informed the approach taken in these fixes.
Keywords: lm studio plugin, web search ai, local web research, fact checking ai, source verification, no api key, offline capable, private browsing ai
A research-grade web search plugin that goes beyond snippets โ it reads pages, verifies claims, detects contradictions, finds primary sources, and enforces cross-source fact verification before asserting anything as true.
Standard search gives you ten links and short snippets. You get:
This plugin fixes all of that.
Load the built plugin in LM Studio.
| Field | Default | Description |
|---|---|---|
| Max Search Results | 8 | Results retrieved per query |
| Max Pages to Read | 3 | Pages actually fetched and read per search |
| Page Fetch Timeout | 8000ms | Per-page timeout before giving up |
| Search Interval | 4000ms | Minimum interval between search requests (all enabled providers). Higher = fewer CAPTCHA blocks, slower. If blocked often, raise to 5000โ6000 ms or use SearXNG. |
| SearXNG | OFF | Self-hosted first tier, opt-in (compact toggle; details on hover). Most users do not host SearXNG โ leave off unless you have a URL. |
| DuckDuckGo | ON | Primary DDG scraper (compact). |
| Bing | ON | Fallback tier (compact). |
OFF | Google scraper, stricter bot protection (compact). | |
| Ecosia | OFF | Bing-powered, privacy (compact). |
| Startpage | ON | Google proxy, privacy (compact). |
| [Dubious] Yandex | OFF | Russian, heavy bot protection (compact). |
| [Dubious] Yahoo | OFF | Bing-powered, ad-heavy (compact). |
| [Dubious] Baidu | OFF |
Before running any search, the plugin calls a clarify step. It detects ambiguity signals in the user's question:
If ambiguity is detected, the LLM asks the user focused questions before searching. This avoids wasted searches and produces a much more targeted answer.
After retrieving search results, the plugin reranks them using embeddings before fetching any pages. This means the pages that actually get read are the most semantically relevant to the query โ not just the top SEO results.
How it works:
nomic-embed-text (requires the model loaded in LM Studio)If the embedding call fails (model not loaded, wrong URL), it falls back silently to the original search engine ranking.
search, search_recent, and search_news each report independent_publishers_read โ the number of distinct root domains among the pages successfully fetched. This feeds a dynamic instruction injected into every result:
| Publishers read | Instruction to the LLM |
|---|---|
| 0 | Do not assert any facts โ re-search or inform the user |
| 1 | Hard UNVERIFIED warning โ call fact_check or label every claim as unverified before presenting |
| 2+ | Report publisher count โ flag any claim supported by only one of them as UNVERIFIED |
This prevents the LLM from repeating a claim as fact just because one website said it.
The system prompt enforces five non-negotiable rules on top of the publisher diversity signal:
Results come from SearXNG (if configured) โ DuckDuckGo โ Startpage (Google proxy, default ON) โ Google (opt-in) โ Ecosia (Bing-powered, opt-in) โ Bing โ [Dubious, opt-in] Yandex / Yahoo / Baidu. No API keys required. All tiers are individually toggleable in settings and share the configurable interval + jitter + randomized User-Agent.
clarify โ Ambiguity check (called automatically first)Always called before any search. Returns either:
STATUS: READY โ question is specific enough, search proceedsSTATUS: CLARIFY โ question is ambiguous, LLM asks user before searchingYou do not need to call this manually. The system prompt enforces it as a mandatory first step.
search โ Core search with page readingThe main tool. Unlike basic search, it fetches and reads the actual page content โ not just snippets.
Returns:
independent_publishers_read โ count of distinct root domains among successful page readsfetch_and_read โ Read a specific URLFetch any URL and return the full readable text content.
Use when:
deep_search โ Multi-angle researchRuns 3โ5 separate searches from different perspectives on the same topic, reads pages for each, and returns everything together. Defeats single-search bias.
Default angles: overview facts, latest research, criticism/limitations, expert consensus.
You can specify your own angles, e.g.:
fact_check โ Verify a specific claimCross-checks a claim across four search angles: direct confirmation, debunking searches, evidence searches, and expert opinion. Returns raw evidence from all angles for the LLM to assess.
Verdict categories: supported, disputed, unsupported, nuanced, uncertain.
verify_statistic โ Verify a number or percentageStatistics are frequently outdated, misquoted, out of context, or fabricated. This tool searches for the stat, its primary source, fact-check results, and updated data.
Example: verify_statistic("90% of startups fail in year one", "venture-backed US tech startups")
find_primary_source โ Trace a claim to its originSecondary sources often distort original findings. This tool searches for the original study, report, official document, or statement where a claim first appeared.
Prioritises: peer-reviewed journals, government reports, official organisation publications over secondary citations.
search_recent โ Time-filtered searchOnly returns results from the specified time window. Prevents stale results from dominating on fast-moving topics.
Windows: day (last 24h), week, month, year.
Returns independent_publishers_read and a dynamic publisher diversity instruction alongside the results.
compare_sources โ Surface agreements and conflictsFetches multiple sources on the same topic and returns them side by side for the LLM to compare framing, spot conflicts, and identify unique claims.
Provide specific URLs to compare, or let it search and pick sources with varied domains automatically.
Returns structured analysis of:
find_expert_views โ Expert consensus and dissentSearches specifically for academic research, official positions, expert interviews, and scientific consensus โ not what random blogs claim experts say.
Covers four angles: expert consensus, peer-reviewed research, official institutional positions, and active scientific debate.
search_academic โ Academic papers onlySearches arXiv, PubMed, and Semantic Scholar for peer-reviewed papers and research publications.
Sources: arxiv, pubmed, semantic_scholar, all.
Fetches paper pages to extract abstracts, methodology, and findings. The LLM is instructed to distinguish preprints from peer-reviewed work, note sample sizes, and not overstate findings.
search_news โ News-specific searchNews-filtered search that actively ranks established journalism above blogs, product pages, and content farms. Runs two queries โ one general, one targeting major news outlets โ then ranks high-credibility results first.
Windows: day, week, month, any.
Unlike search_recent (which filters by date), this filters by source type โ it's about journalistic sourcing, not just recency. Best for: breaking news, corporate announcements, policy changes, anything where "who is reporting it" matters.
Returns independent_publishers_read and a dynamic publisher diversity instruction alongside the results.
research_topic โ Full multi-step research briefRuns multiple searches from different angles, reads key pages, and instructs the LLM to produce a structured research brief: overview, established facts, contested areas, expert consensus, open questions, key sources, and confidence assessment.
Depths:
overview โ 3 angles, 2 pages eachdetailed โ 5 angles, 2 pages each (default)comprehensive โ 7 angles, 3 pages eachcheck_source โ Source credibility assessmentAssesses a URL or domain and returns its credibility type, known signals, reputation search results, and red flags to watch for.
Domain types: government, academic institution, academic/research platform, established news outlet, encyclopedia, user-generated content, unknown.
Credibility levels: high, medium, low, unknown.
Red flags checked:
Every search result and fetched page gets a credibility assessment based on domain signals:
| Domain Type | Credibility | Examples |
|---|---|---|
| Government | HIGH | .gov, .mil, WHO, CDC |
| Academic institution | HIGH | .edu, .ac.uk, universities |
| Academic platforms | HIGH | arXiv, PubMed, Semantic Scholar |
| Established news | HIGH | Reuters, AP, BBC, Nature, NYT |
| Wikipedia | MEDIUM | Good overview, verify citations |
| User-generated / blogs | LOW | Blogspot, WordPress, Reddit, Quora |
| Unknown | UNKNOWN | Check About page and author credentials |
The plugin's system prompt instructs the LLM to:
Simple fact:
"What is the Dunning-Kruger effect?" โ
clarify(READY) โsearch
Ambiguous query:
"Tell me about python" โ
clarify(CLARIFY) โ asks: "Do you mean the programming language or the snake?" โsearch
Verify a claim:
"Is it true that we only use 10% of our brains?" โ
clarify(READY) โfact_check
Verify a statistic:
"Someone told me 50,000 species go extinct every year. Is that right?" โ
clarify(READY) โverify_statistic
Recent developments:
"What's happened with GPT-5 in the last week?" โ
clarify(READY) โsearch_recent(window: "week")
Compare perspectives:
"What do different sources say about seed oils and health?" โ
clarify(READY) โcompare_sourcesordeep_search
Scientific consensus:
"What does the research actually say about intermittent fasting?" โ
clarify(READY) โfind_expert_views+search_academic
Deep research:
"Give me a thorough research brief on quantum error correction" โ
clarify(READY) โresearch_topic(depth: "comprehensive")
Read a specific article:
"Can you read this paper and summarise the key findings? [url]" โ
clarify(READY) โfetch_and_read
Check if a source is reliable:
"Is naturalhealth365.com a reliable source?" โ
clarify(READY) โcheck_source
AI capability claim:
"I read that [model X] achieves 95% accuracy on [benchmark]. Is that right?" โ
clarify(READY) โfact_checkโ vendor blogs and press releases are not accepted as evidence; requires independent academic or journalistic confirmation
Other altra plugins that include web search as a secondary capability will automatically defer to this plugin when it is installed alongside them. Their duplicate search/fetch tools are omitted at startup so this plugin's richer versions take over:
| Plugin | Tools deferred to web-search |
|---|---|
altra/research | search_sources, read_source |
altra/ideas | research |
altra/high-perf-tools | fetch_url, search_web |
Installing this plugin is the recommended way to get the best search quality across all plugins at once. Each affected plugin also shows a tip on first message when this plugin is not installed.
The ddgSearch() function had multiple flaws that caused search results to silently return empty after one or two successful calls per session:
No TLS fingerprint spoofing โ DuckDuckGo performs TLS fingerprinting to detect bots. The native call lacks a realistic TLS fingerprint, so DDG identifies it as a bot and serves a CAPTCHA challenge page ("Select all squares containing a duck") instead of search results. This is the primary cause of persistent empty results.
got-scraping for anti-bot evasion (new dependency):
All search backend requests and page fetches now use got-scraping with headerGeneratorOptions (Chrome 120+, desktop, Linux, en-US). This generates realistic browser headers and TLS fingerprints that bypass DDG's bot detection. Native fetch() is used as a fallback only.
CAPTCHA detection (isDuckDuckGoCaptcha):
Detects DDG's bot challenge pages by checking for anomaly-modal, cc=botnet, challenge-form, and related markers. When detected, throws a clear error instead of silently returning empty results.
Rate limiter (makeRateLimiter + searchRateLimiter):
Enforces a 2000ms minimum interval between search backend requests, preventing DuckDuckGo from blocking or rate-limiting the plugin.
Retry logic (DDG_MAX_RETRIES = 3, DDG_RETRY_WAIT_MS = 2500):
Retries DuckDuckGo up to 3 times with 2.5s waits when results are empty. CAPTCHA errors and network failures fall through immediately to Bing without retrying.
Diagnostic reporting (SearchDiagnostic interface):
Every search call now returns a search_diagnostics array alongside results. Each entry tracks:
backend โ which search engine was tried (searxng, ddg, bing)ok โ whether it returned resultsAll 13 tool implementations updated:
Each tool now includes search_diagnostics in its JSON response, so the model and user can see exactly which backend succeeded or failed and why.
When searches fail, instead of silently empty results, the response includes diagnostics:
This tells you exactly which backends were tried, how many attempts were made, and what went wrong โ so you and the model can diagnose and adapt.
search_recent over-triggering (model defaulted to time-filtered search)Previously the model invoked search_recent/search_news too often even for evergreen queries. Three guardrails were added:
Category split (config.ts:33-68, toolsProvider.ts:92-100,1067-1085): Search providers are now grouped into Regular (SearXNG, DDG, Bing, Google, Ecosia, Startpage) and [Dubious] (Yandex, Baidu, Yahoo). Dubious engines are those with heavier bias/censorship, ad-heavy or less reliable indexes, and aggressive anti-scraping. Each has an individual boolean toggle in the GUI (enableYandex / enableYahoo / enableBaidu โ all OFF by default, opt-in only) with subtitle warnings. Regular toggles are enableEcosia (OFF by default) and enableStartpage (ON by default) in addition to the existing enableSearxng/enableDDG/enableBing/enableGoogle. The unified chain is now ; disabled tiers emit diagnostics. Documents the new toggles; documents the full order.
All new scrapers reuse the hardened fetchSearchHtml() (TLS spoof, jitter interval, randomized UA from userAgents, proper Referer) and report backend: "ecosia"|"startpage"|"yandex"|"yahoo"|"baidu" diagnostics with CAPTCHA/empty handling consistent with DDG/Google/Bing.
promptPreprocessor.ts:74-81, toolsProvider.ts:39-62,173-195,2170-2245Squashed several subtle loops that could still trap the model even after the earlier re-search fix:
Thanks to Altra for making something cool, and giving it to the world to toy around!
Additional "Thank you!" to brius and nub235 for field expertise; The engineering solutions in your code served as a reference implementation for fixes here. Your use of got-scraping for anti-bot evasion, structured error reporting, and rate limiting informed the approach taken in these fixes.
| Chinese, censored, aggressive (compact). |
| Custom User Agents | 4 defaults | Paragraph field โ one per line, randomly chosen per request. Long strings wrap and you can scroll/drag/arrow-key inside the field. Prefix a line with # or 0:: to disable without deleting; remove prefix or use 1:: to re-enable. Delete a line to remove entirely. Leave empty for built-in randomized header generator. (Previous stringArray UI had no per-item toggle and truncated long UAs โ this paragraph field fixes both; SDK limitation: stringArray cannot have per-item enable/disable toggles, so prefix is the workaround.) |
| Search Language | en-us | Language/region for results |
| SearXNG URL | (blank) | Recommended. Self-hosted SearXNG instance. Falls back to DDG โ Startpage โ Google โ Ecosia โ Bing โ (Dubious) Yandex/Yahoo/Baidu if blank. DDG/Bing/Google etc. may block headless requests โ SearXNG is the reliable path. |
| Search Recency Window | any | Limit general search to: day, week, month, year, or any (no filter, default). Use search_recent / search_news when you explicitly need recent results. |
| LM Studio URL | http://localhost:1234 | Used for embedding-based result reranking via nomic-embed-text |
fact_check or compare_sources.verify_statistic before asserting the number.clarify first โ ask focused questions before searching if the query is ambiguousfetch()Silent error swallowing โ all catch { /* fall through */ } blocks discarded errors completely. When DuckDuckGo blocked or rate-limited requests, there was zero indication of what went wrong.
No rate limiting โ only 300ms sleeps between requests. DuckDuckGo silently returns empty results when hit too fast.
No retry logic โ if DDG returned empty (due to blocking), it immediately fell through to Bing (which often also fails) without retrying.
resultCounterror โ error message if it failedattempts โ how many retries were attempted (DDG only)promptPreprocessor.ts): TOOL SELECTION GUIDE now marks search as DEFAULT โ 90% of queries and restricts search_recent/search_news to RECENCY-SENSITIVE ONLY โ only when the user explicitly says latest/recent/current/breaking/today/this weekโฆ or the query contains time-sensitive keywords (latest, recent, current, today, breaking, news, 2025/2026, this week/month) or the topic is inherently time-sensitive (stock price, weather, live scores, breaking news). Otherwise: When in doubt, use 'search'.toolsProvider.ts): search is now DEFAULT for 90% of queries / FIRST CHOICE; search_recent and search_news both start with DO NOT use by default and Use ONLY when (a) explicit user request / (b) explicit time keywords / (c) inherently time-sensitive topic.config.ts/config.js): searchRecencyWindow default changed year โ any (no filter). This makes general search timeless and prevents evergreen topics from being silently filtered; explicit recency now requires search_recent/search_news with a window.Removed re-search loop instruction (toolsProvider.ts:106-118): publisherDiversityInstruction() previously told the model re-search or inform the user when no pages were read. This caused the model to retry the exact same query identically after a DDG CAPTCHA/block, creating an infinite loop. Replaced with: check search_diagnostics for CAPTCHA/bot errors, do not immediately retry the same query, you may wait a few seconds and retry once (optionally rephrased), and if this is the 2nd/3rd consecutive block, stop and clearly inform the user (explain the engine is temporarily blocking automated requests, list search_diagnostics, suggest alternatives). A new helper botBlockGuidance() appends the same guidance to every tool returning search_diagnostics (search, search_recent, search_news, deep_search, fact_check, verify_statistic, find_primary_source, compare_sources, find_expert_views, search_academic, research_topic).
Inspected bot-protection countermeasures (toolsProvider.ts:366-411): DDG blocks were frequent because the fingerprint was too static and the cadence too regular.
fetchSearchHtml() and fetchPage() now use diversified headerGeneratorOptions โ chrome 120โ130 and firefox 115โ128, devices: [desktop], locales: [en-US, en], operatingSystems: [windows, macos, linux] instead of fixed chrome 120 / linux / en-US.Referer: https://duckduckgo.com/ to search fetches so requests look like normal navigation.isDuckDuckGoCaptcha() from 5 to 12+ signals (case-insensitive captcha, please verify you are a human, verifying you are human, temporarily blocked, rate limit, automated requests, enable javascript+continue, challenges.cloudflare.com, etc.) so blocks are correctly classified instead of appearing as empty results.DDG_RETRY_WAIT_MS 2500 โ 3500 ms (plus rate-limiter jitter) to avoid hammering a blocked endpoint.Increased interval and made it configurable (config.ts:22-31, toolsProvider.ts:72-84,673-699):
2000 ms โ 4000 ms (was the brius-web-search 2 s value). Higher values reduce CAPTCHA frequency; multi-angle searches are slower but far more reliable.searchIntervalMs (numeric 1000โ10000 ms, step 500, slider, default 4000) with subtitle explaining the trade-off and SearXNG alternative. The plugin syncs configuredSearchIntervalMs from cfg.get("searchIntervalMs") before every ddg()/sar() call.makeRateLimiter() now takes () => number and adds 0โ600 ms jitter on every wait, so requests are not perfectly periodic (a known bot signal).Google search support (toolsProvider.ts:557-623, 625-631, 644-730): Google sometimes yields better results than DDG/Bing. Added isGoogleCaptcha() (detects Our systems have detected unusual traffic, /sorry/index, recaptcha + google.com/sorry, unusual traffic + captcha, etc.) and scrapeGoogle() (builds https://www.google.com/search?q=โฆ&num=&hl=en&gl=us&pws=0&tbs=qdr:X when a time window is requested, reuses the hardened fetchSearchHtml() with TLS spoofing + custom UA, parses yuRUbf blocks and /url?q= redirects with a generic <a><h3> fallback). Integrated into the unified fallback chain as Tier 3 between DDG and Bing: SearXNG โ DDG โ Google โ Bing. Google respects the same searchRateLimiter() + jitter and contributes backend: "google" entries to search_diagnostics.
Provider enable/disable GUI (config.ts:33-48, toolsProvider.ts:89-102,645-730,848-872): Added enableSearxng (default ON), enableDDG (ON), enableBing (ON), enableGoogle (OFF by default) as boolean fields in LM Studio's plugin settings. The plugin syncs configuredProviders from cfg before every ddg()/sar() call; disabled providers emit skipped โ disabled in config diagnostics and are bypassed. This lets you run DDG-only, Google+Bing, SearXNG-only, etc., without editing code. Order is fixed (SearXNG โ DDG โ Google โ Bing) but each tier can be individually disabled.
Custom User-Agent multi-select (config.ts:50-64, toolsProvider.ts:104-131,348-365,381-393,422-472,629-631,996-998): Added userAgents stringArray with 4 sensible defaults (Chrome Win/Mac, Firefox Linux, Safari Mac). GUI shows them as chips โ add via input, delete via ร. Per-item enable/disable without deleting: prefix an entry with # or 0:: (or disabled::) to disable it, remove the prefix or use 1:: to re-enable; empty or non-string entries are ignored. The plugin filters to enabled entries via getEnabledUserAgents(), and pickRandomUserAgent() chooses one at random per request (search and page fetch). Both fetchSearchHtml() (search backends) and fetchPage() (article reads + SearXNG) now send the chosen UA as User-Agent (overriding the header generator) and set a proper Referer (<origin>/ from the request URL). Native fetch fallbacks also use the same random UA. If the list is empty or all entries are disabled, the built-in randomized got-scraping header generator is used as before.
skipped โ disabled in configREADME.md:54-62README.md:120Ecosia (toolsProvider.ts:632-670) โ regular, disabled by default: Bing-powered, privacy-focused. Scrapes https://www.ecosia.org/search?q=. Useful as a Bing alternative if Bing blocks (different domain, same Bing index). Parser looks for <article class="result"> / <div class="result"> with <a href> + <h2> title and <p> snippet, with generic <a><h2/3> fallback. isEcosiaCaptcha() checks captcha+ecosia, verify you are human, unusual traffic, challenges.cloudflare.com.
Startpage (toolsProvider.ts:672-722) โ regular, enabled by default: Google proxy, privacy-focused. Scrapes https://www.startpage.com/sp/search?query=. Provides Google results without tracking; if Google's sorry/index CAPTCHA blocks, Startpage via a different frontend may still succeed. Note on difficulty: Startpage is Cloudflare-protected (challenges.cloudflare.com / cf-challenge / DDoS protection by Cloudflare / Please turn JavaScript on), so scraping is inherently less reliable than DDG/Bing. Implementation tries w-gl__result blocks with w-gl__description snippets, but will often fall through to the next tier; diagnostics will show Startpage exception: โฆCloudflare/CAPTCHAโฆ rather than silently empty.
Yandex (toolsProvider.ts:724-767) โ dubious, OFF: Scrapes https://yandex.com/search?text=. isYandexCaptcha() detects SmartCaptcha/showcaptcha/please confirm you are not a robot/captcha+yandex. Parser looks for li.serp-item + <a href> and OrganicText snippet, with generic fallback. Yandex heavily JS-rendered and SmartCaptcha is frequent; success rate may be low without SearXNG.
Yahoo (toolsProvider.ts:769-810) โ dubious, OFF: Scrapes https://search.yahoo.com/search?p=. Yahoo is Bing-powered and ad-heavy. isYahooCaptcha() checks captcha+yahoo/verify you are human. Parser for div.algo / dd algo with h3>a and compText snippet. Yahoo wraps result URLs in r.search.yahoo.com redirects โ those are skipped.
Baidu (toolsProvider.ts:812-863) โ dubious, OFF: Scrapes https://www.baidu.com/s?wd=. Most difficult: Baidu uses c-container blocks, baidu.com/link?url= redirects that require decoding, Chinese ๅฎๅ
จ้ช่ฏ (security verification) challenge, and GBK encoding quirks. Parser handles c-container โ h3>a + c-abstract snippet, decodes url= param, falls back to generic <a> scan. Expect low success outside Chinese queries; leave disabled unless you specifically need Baidu coverage.
safe_impl hint now branches on bot-block (toolsProvider.ts:53-62): previously always Read the error above, adjust the parameters if needed, and retry. โ now detects captcha|bot challenge|blocked|botnet|challenge-form|unusual traffic|Cloudflare|SmartCaptcha|429|403 and returns an anti-loop hint (Do NOT retry same query immediatelyโฆ wait, rephrase, try different enabled provider, or inform userโฆ max 2 identical calls per turn). Prevents hint contradicting botBlockGuidance.publisherDiversityInstruction:0-pages hardened (toolsProvider.ts:173-188): now also handles skipped โ disabled in config (tell user to enable a provider instead of retrying) and genuine 0 results, no CAPTCHA (rephrase broader or inform no results, do not loop). Expanded provider list to SearXNG/DDG/Startpage/Google/Ecosia/Bing + dubious in suggestions.toolsProvider.ts:191-194): WARNING: All pages read are from same publisherโฆ Call fact_check on key claims โ now Consider calling fact_check ONLY if claim is high-stakes and you have not already called fact_check for this claim; otherwise label unverified. Do not call fact_check repeatedly โ max once per claim per turn. Prevents search 1 publisher โ fact_check (4ร searches) โ also single-source โ fact_check โฆ amplification.check_source rate-limited (toolsProvider.ts:2238-2245): instruction now Limit use: do not call check_source for every URL โ max 1โ2 most important domains per turn. If already checked, reuse prior verdict. + botBlockGuidance(allDiags) attached. Prevents model calling check_source for each of 3 results โ 6 extra searches.LOOP PREVENTION (hard) in SYSTEM_RULES (promptPreprocessor.ts:74-81): new top-level rule NEVER call same tool with identical parameters more than twice per turn. After 2 failures (empty, CAPTCHA, tool_error) STOP and summarize search_diagnosticsโฆ suggest wait 30โ60s / enable different provider / rephrase broader / provide URL. If publisherDiversityInstruction / botBlockGuidance / safe_impl hint says STOP, obey immediately. Max one fact_check/verify_statistic/check_source per claim/domain per turn, max one deep_search/research_topic per topic per turn.cd web-search-plugin
npm install
npx tsc
search(query, max_pages_to_read?)
fetch_and_read(url, max_chars?)
deep_search(topic, angles?, pages_per_angle?)
angles: ["economic impact", "environmental cost", "industry response", "regulatory landscape"]
fact_check(claim)
verify_statistic(statistic, context?)
find_primary_source(claim, domain?)
search_recent(query, window?, read_pages?)
compare_sources(topic, urls?, num_sources?)
find_expert_views(topic, field?)
search_academic(topic, source?, year_from?)
search_news(query, window?, read_pages?)
research_topic(topic, depth?, focus?)
check_source(url)
{
"search_diagnostics": [
{ "backend": "ddg", "ok": false, "resultCount": 0,
"error": "DDG returned a bot challenge (CAPTCHA) โ requests are being blocked by DuckDuckGo's anti-bot system",
"attempts": 1 },
{ "backend": "bing", "ok": true, "resultCount": 5 }
]
}
| Chinese, censored, aggressive (compact). |
| Custom User Agents | 4 defaults | Paragraph field โ one per line, randomly chosen per request. Long strings wrap and you can scroll/drag/arrow-key inside the field. Prefix a line with # or 0:: to disable without deleting; remove prefix or use 1:: to re-enable. Delete a line to remove entirely. Leave empty for built-in randomized header generator. (Previous stringArray UI had no per-item toggle and truncated long UAs โ this paragraph field fixes both; SDK limitation: stringArray cannot have per-item enable/disable toggles, so prefix is the workaround.) |
| Search Language | en-us | Language/region for results |
| SearXNG URL | (blank) | Recommended. Self-hosted SearXNG instance. Falls back to DDG โ Startpage โ Google โ Ecosia โ Bing โ (Dubious) Yandex/Yahoo/Baidu if blank. DDG/Bing/Google etc. may block headless requests โ SearXNG is the reliable path. |
| Search Recency Window | any | Limit general search to: day, week, month, year, or any (no filter, default). Use search_recent / search_news when you explicitly need recent results. |
| LM Studio URL | http://localhost:1234 | Used for embedding-based result reranking via nomic-embed-text |
fact_check or compare_sources.verify_statistic before asserting the number.clarify first โ ask focused questions before searching if the query is ambiguousfetch()Silent error swallowing โ all catch { /* fall through */ } blocks discarded errors completely. When DuckDuckGo blocked or rate-limited requests, there was zero indication of what went wrong.
No rate limiting โ only 300ms sleeps between requests. DuckDuckGo silently returns empty results when hit too fast.
No retry logic โ if DDG returned empty (due to blocking), it immediately fell through to Bing (which often also fails) without retrying.
resultCounterror โ error message if it failedattempts โ how many retries were attempted (DDG only)promptPreprocessor.ts): TOOL SELECTION GUIDE now marks search as DEFAULT โ 90% of queries and restricts search_recent/search_news to RECENCY-SENSITIVE ONLY โ only when the user explicitly says latest/recent/current/breaking/today/this weekโฆ or the query contains time-sensitive keywords (latest, recent, current, today, breaking, news, 2025/2026, this week/month) or the topic is inherently time-sensitive (stock price, weather, live scores, breaking news). Otherwise: When in doubt, use 'search'.toolsProvider.ts): search is now DEFAULT for 90% of queries / FIRST CHOICE; search_recent and search_news both start with DO NOT use by default and Use ONLY when (a) explicit user request / (b) explicit time keywords / (c) inherently time-sensitive topic.config.ts/config.js): searchRecencyWindow default changed year โ any (no filter). This makes general search timeless and prevents evergreen topics from being silently filtered; explicit recency now requires search_recent/search_news with a window.Removed re-search loop instruction (toolsProvider.ts:106-118): publisherDiversityInstruction() previously told the model re-search or inform the user when no pages were read. This caused the model to retry the exact same query identically after a DDG CAPTCHA/block, creating an infinite loop. Replaced with: check search_diagnostics for CAPTCHA/bot errors, do not immediately retry the same query, you may wait a few seconds and retry once (optionally rephrased), and if this is the 2nd/3rd consecutive block, stop and clearly inform the user (explain the engine is temporarily blocking automated requests, list search_diagnostics, suggest alternatives). A new helper botBlockGuidance() appends the same guidance to every tool returning search_diagnostics (search, search_recent, search_news, deep_search, fact_check, verify_statistic, find_primary_source, compare_sources, find_expert_views, search_academic, research_topic).
Inspected bot-protection countermeasures (toolsProvider.ts:366-411): DDG blocks were frequent because the fingerprint was too static and the cadence too regular.
fetchSearchHtml() and fetchPage() now use diversified headerGeneratorOptions โ chrome 120โ130 and firefox 115โ128, devices: [desktop], locales: [en-US, en], operatingSystems: [windows, macos, linux] instead of fixed chrome 120 / linux / en-US.Referer: https://duckduckgo.com/ to search fetches so requests look like normal navigation.isDuckDuckGoCaptcha() from 5 to 12+ signals (case-insensitive captcha, please verify you are a human, verifying you are human, temporarily blocked, rate limit, automated requests, enable javascript+continue, challenges.cloudflare.com, etc.) so blocks are correctly classified instead of appearing as empty results.DDG_RETRY_WAIT_MS 2500 โ 3500 ms (plus rate-limiter jitter) to avoid hammering a blocked endpoint.Increased interval and made it configurable (config.ts:22-31, toolsProvider.ts:72-84,673-699):
2000 ms โ 4000 ms (was the brius-web-search 2 s value). Higher values reduce CAPTCHA frequency; multi-angle searches are slower but far more reliable.searchIntervalMs (numeric 1000โ10000 ms, step 500, slider, default 4000) with subtitle explaining the trade-off and SearXNG alternative. The plugin syncs configuredSearchIntervalMs from cfg.get("searchIntervalMs") before every ddg()/sar() call.makeRateLimiter() now takes () => number and adds 0โ600 ms jitter on every wait, so requests are not perfectly periodic (a known bot signal).Google search support (toolsProvider.ts:557-623, 625-631, 644-730): Google sometimes yields better results than DDG/Bing. Added isGoogleCaptcha() (detects Our systems have detected unusual traffic, /sorry/index, recaptcha + google.com/sorry, unusual traffic + captcha, etc.) and scrapeGoogle() (builds https://www.google.com/search?q=โฆ&num=&hl=en&gl=us&pws=0&tbs=qdr:X when a time window is requested, reuses the hardened fetchSearchHtml() with TLS spoofing + custom UA, parses yuRUbf blocks and /url?q= redirects with a generic <a><h3> fallback). Integrated into the unified fallback chain as Tier 3 between DDG and Bing: SearXNG โ DDG โ Google โ Bing. Google respects the same searchRateLimiter() + jitter and contributes backend: "google" entries to search_diagnostics.
Provider enable/disable GUI (config.ts:33-48, toolsProvider.ts:89-102,645-730,848-872): Added enableSearxng (default ON), enableDDG (ON), enableBing (ON), enableGoogle (OFF by default) as boolean fields in LM Studio's plugin settings. The plugin syncs configuredProviders from cfg before every ddg()/sar() call; disabled providers emit skipped โ disabled in config diagnostics and are bypassed. This lets you run DDG-only, Google+Bing, SearXNG-only, etc., without editing code. Order is fixed (SearXNG โ DDG โ Google โ Bing) but each tier can be individually disabled.
Custom User-Agent multi-select (config.ts:50-64, toolsProvider.ts:104-131,348-365,381-393,422-472,629-631,996-998): Added userAgents stringArray with 4 sensible defaults (Chrome Win/Mac, Firefox Linux, Safari Mac). GUI shows them as chips โ add via input, delete via ร. Per-item enable/disable without deleting: prefix an entry with # or 0:: (or disabled::) to disable it, remove the prefix or use 1:: to re-enable; empty or non-string entries are ignored. The plugin filters to enabled entries via getEnabledUserAgents(), and pickRandomUserAgent() chooses one at random per request (search and page fetch). Both fetchSearchHtml() (search backends) and fetchPage() (article reads + SearXNG) now send the chosen UA as User-Agent (overriding the header generator) and set a proper Referer (<origin>/ from the request URL). Native fetch fallbacks also use the same random UA. If the list is empty or all entries are disabled, the built-in randomized got-scraping header generator is used as before.
skipped โ disabled in configREADME.md:54-62README.md:120Ecosia (toolsProvider.ts:632-670) โ regular, disabled by default: Bing-powered, privacy-focused. Scrapes https://www.ecosia.org/search?q=. Useful as a Bing alternative if Bing blocks (different domain, same Bing index). Parser looks for <article class="result"> / <div class="result"> with <a href> + <h2> title and <p> snippet, with generic <a><h2/3> fallback. isEcosiaCaptcha() checks captcha+ecosia, verify you are human, unusual traffic, challenges.cloudflare.com.
Startpage (toolsProvider.ts:672-722) โ regular, enabled by default: Google proxy, privacy-focused. Scrapes https://www.startpage.com/sp/search?query=. Provides Google results without tracking; if Google's sorry/index CAPTCHA blocks, Startpage via a different frontend may still succeed. Note on difficulty: Startpage is Cloudflare-protected (challenges.cloudflare.com / cf-challenge / DDoS protection by Cloudflare / Please turn JavaScript on), so scraping is inherently less reliable than DDG/Bing. Implementation tries w-gl__result blocks with w-gl__description snippets, but will often fall through to the next tier; diagnostics will show Startpage exception: โฆCloudflare/CAPTCHAโฆ rather than silently empty.
Yandex (toolsProvider.ts:724-767) โ dubious, OFF: Scrapes https://yandex.com/search?text=. isYandexCaptcha() detects SmartCaptcha/showcaptcha/please confirm you are not a robot/captcha+yandex. Parser looks for li.serp-item + <a href> and OrganicText snippet, with generic fallback. Yandex heavily JS-rendered and SmartCaptcha is frequent; success rate may be low without SearXNG.
Yahoo (toolsProvider.ts:769-810) โ dubious, OFF: Scrapes https://search.yahoo.com/search?p=. Yahoo is Bing-powered and ad-heavy. isYahooCaptcha() checks captcha+yahoo/verify you are human. Parser for div.algo / dd algo with h3>a and compText snippet. Yahoo wraps result URLs in r.search.yahoo.com redirects โ those are skipped.
Baidu (toolsProvider.ts:812-863) โ dubious, OFF: Scrapes https://www.baidu.com/s?wd=. Most difficult: Baidu uses c-container blocks, baidu.com/link?url= redirects that require decoding, Chinese ๅฎๅ
จ้ช่ฏ (security verification) challenge, and GBK encoding quirks. Parser handles c-container โ h3>a + c-abstract snippet, decodes url= param, falls back to generic <a> scan. Expect low success outside Chinese queries; leave disabled unless you specifically need Baidu coverage.
safe_impl hint now branches on bot-block (toolsProvider.ts:53-62): previously always Read the error above, adjust the parameters if needed, and retry. โ now detects captcha|bot challenge|blocked|botnet|challenge-form|unusual traffic|Cloudflare|SmartCaptcha|429|403 and returns an anti-loop hint (Do NOT retry same query immediatelyโฆ wait, rephrase, try different enabled provider, or inform userโฆ max 2 identical calls per turn). Prevents hint contradicting botBlockGuidance.publisherDiversityInstruction:0-pages hardened (toolsProvider.ts:173-188): now also handles skipped โ disabled in config (tell user to enable a provider instead of retrying) and genuine 0 results, no CAPTCHA (rephrase broader or inform no results, do not loop). Expanded provider list to SearXNG/DDG/Startpage/Google/Ecosia/Bing + dubious in suggestions.toolsProvider.ts:191-194): WARNING: All pages read are from same publisherโฆ Call fact_check on key claims โ now Consider calling fact_check ONLY if claim is high-stakes and you have not already called fact_check for this claim; otherwise label unverified. Do not call fact_check repeatedly โ max once per claim per turn. Prevents search 1 publisher โ fact_check (4ร searches) โ also single-source โ fact_check โฆ amplification.check_source rate-limited (toolsProvider.ts:2238-2245): instruction now Limit use: do not call check_source for every URL โ max 1โ2 most important domains per turn. If already checked, reuse prior verdict. + botBlockGuidance(allDiags) attached. Prevents model calling check_source for each of 3 results โ 6 extra searches.LOOP PREVENTION (hard) in SYSTEM_RULES (promptPreprocessor.ts:74-81): new top-level rule NEVER call same tool with identical parameters more than twice per turn. After 2 failures (empty, CAPTCHA, tool_error) STOP and summarize search_diagnosticsโฆ suggest wait 30โ60s / enable different provider / rephrase broader / provide URL. If publisherDiversityInstruction / botBlockGuidance / safe_impl hint says STOP, obey immediately. Max one fact_check/verify_statistic/check_source per claim/domain per turn, max one deep_search/research_topic per topic per turn.cd web-search-plugin
npm install
npx tsc
search(query, max_pages_to_read?)
fetch_and_read(url, max_chars?)
deep_search(topic, angles?, pages_per_angle?)
angles: ["economic impact", "environmental cost", "industry response", "regulatory landscape"]
fact_check(claim)
verify_statistic(statistic, context?)
find_primary_source(claim, domain?)
search_recent(query, window?, read_pages?)
compare_sources(topic, urls?, num_sources?)
find_expert_views(topic, field?)
search_academic(topic, source?, year_from?)
search_news(query, window?, read_pages?)
research_topic(topic, depth?, focus?)
check_source(url)
{
"search_diagnostics": [
{ "backend": "ddg", "ok": false, "resultCount": 0,
"error": "DDG returned a bot challenge (CAPTCHA) โ requests are being blocked by DuckDuckGo's anti-bot system",
"attempts": 1 },
{ "backend": "bing", "ok": true, "resultCount": 5 }
]
}