Project Files
README
📡 RSS RAG Cache Plugin for LM Studio
Turn LM Studio into your own private, intelligent news archive. This plugin fetches your favorite RSS feeds, chunks the articles, embeds them using your local LM Studio server, and stores them in a local SQLite database for Retrieval-Augmented Generation (RAG).
Instead of searching the web, your LLM can instantly query your local cache of recent articles, blog posts, and newsletters.
✨ Features
Retry-After (capped at 15s), so flaky or rate-limited feeds fail gracefully instead of erroring out.news, article, post, ...).node:sqlite module (WAL mode). No C++ addons or Visual Studio build tools required!feeds.txt and cache.db in ~/.lmstudio/plugin-data/rss-rag-cache/ so they survive plugin updates.📋 Requirements
nomic-embed-text-v1.5) and start the local server on port 1234.🚀 Installation
feeds.txt in your plugin-data folder on the first run.⚙️ Configuration
1. Adding Feeds (feeds.txt)
Your feeds file is located at:
C:\Users\YOUR_USERNAME\.lmstudio\plugin-data\rss-rag-cache\feeds.txt~/.lmstudio/plugin-data/rss-rag-cache/feeds.txtAdd one RSS/Atom URL per line. You can also add standard webpage URLs (the plugin will scrape them automatically). Lines starting with # are ignored.
https://news.ycombinator.com/rss https://www.theverge.com/rss/index.xml # Non-RSS webpages (scraped automatically, slower): https://example.com/news
2. Plugin Settings UI
You can adjust these values in the LM Studio plugin settings panel:
| Setting | Default | Description |
|---|---|---|
retentionDays | 90 | How long to keep articles (30, 60, 90, or 365 days). |
refreshIntervalHours | 6 | How often the plugin automatically fetches new RSS items in the background. |
embeddingModel | nomic-embed-text-v1.5 | The exact model ID loaded in your LM Studio server. |
topK | 8 | How many unique articles to return per search. |
chunkSize | 800 | Length (in characters) of each text chunk used for embedding and search. |
embedBatchSize | 16 | How many chunks are sent to the embedding model per request (higher = faster but more VRAM). |
maxConcurrentFeeds | 8 | How many RSS/Atom feeds are fetched at the same time. |
htmlLinkPatterns | news,article,post,blog,stories,press | Comma-separated URL keywords used when scraping non-RSS pages. |
scrapeFullText | Off | For HTML-scraped pages, also download each article's full body text (slower refresh, richer content). |
retentionBasis | publish_date | Whether old entries are removed by publish date or cache date. |
refreshOnStartup | On | Automatically fetch feeds when the plugin loads. |
🛠️ LM Studio Tools
The plugin exposes these tools to the LLM:
refresh_rss_feeds — Force-refreshes all feeds listed in feeds.txt, indexes new items, generates embeddings, and applies the retention policy. Returns a report of successes, failures, and added articles.search_rss_cache — Hybrid (keyword + semantic) search, optionally filtered by days, feed (URL substring), or author.list_rss_feeds — Lists all configured feed URLs plus cache stats: entry/chunk counts, missing embeddings, DB size, and last refresh report.get_recent_cache_articles — Returns the most recent cached articles. Use this to see what news is available without guessing keywords.add_rss_feed — Adds an RSS/Atom feed URL or webpage URL to feeds.txt.remove_rss_feed — Removes a feed URL from feeds.txt (cached articles kept until pruned).test_feed — Checks whether a URL is a working RSS/Atom feed or scrapeable HTML page, and how many items it exposes.prune_removed_feeds — Deletes cached articles whose source feed is no longer in feeds.txt.reindex_missing_embeddings — Re-embeds chunks stored as text-only while the embedding model was offline.🧠 How It Works
fast-xml-parser (with a legacy regex fallback). If a URL returns HTML, it uses Cheerio to scrape the page for article links.maxConcurrentFeeds). HTML pages are fetched one-by-one with a 1.5–3 second delay to avoid anti-bot firewalls. Requests retry with backoff on 429/403/5xx./v1/embeddings endpoint in batches. If the model is offline, chunks are stored text-only and reindex_missing_embeddings completes them later.cache.db). Newer items list first; retention purges by publish or cache date.🐛 Troubleshooting
Search returns 0 matches
Ensure your embedding model is loaded and the server is running. If the model is offline, the plugin falls back to keyword search — but if your articles don't contain the exact keywords, they won't be found. Run list_rss_feeds to check that chunks > 0, or reindex_missing_embeddings once the model is loaded.
HTTP 403 Forbidden on a feed Some websites block automated requests. The plugin uses a Chrome User-Agent, throttles HTML requests, and retries with backoff, but strict firewalls might still block it. Try an alternative RSS URL or an RSS Bridge service.
HTTP 429 Too Many Requests
Sites like Reddit aggressively block Node.js fetch requests. If this happens, use an RSS Bridge URL (e.g., https://rss-bridge.org/bridge01/?action=display&bridge=RedditBridge&format=Atom) instead of the native .rss link.
Project Files
README
📡 RSS RAG Cache Plugin for LM Studio
Turn LM Studio into your own private, intelligent news archive. This plugin fetches your favorite RSS feeds, chunks the articles, embeds them using your local LM Studio server, and stores them in a local SQLite database for Retrieval-Augmented Generation (RAG).
Instead of searching the web, your LLM can instantly query your local cache of recent articles, blog posts, and newsletters.
✨ Features
Retry-After (capped at 15s), so flaky or rate-limited feeds fail gracefully instead of erroring out.news, article, post, ...).node:sqlite module (WAL mode). No C++ addons or Visual Studio build tools required!feeds.txt and cache.db in ~/.lmstudio/plugin-data/rss-rag-cache/ so they survive plugin updates.📋 Requirements
nomic-embed-text-v1.5) and start the local server on port 1234.🚀 Installation
feeds.txt in your plugin-data folder on the first run.⚙️ Configuration
1. Adding Feeds (feeds.txt)
Your feeds file is located at:
C:\Users\YOUR_USERNAME\.lmstudio\plugin-data\rss-rag-cache\feeds.txt~/.lmstudio/plugin-data/rss-rag-cache/feeds.txtAdd one RSS/Atom URL per line. You can also add standard webpage URLs (the plugin will scrape them automatically). Lines starting with # are ignored.
https://news.ycombinator.com/rss https://www.theverge.com/rss/index.xml # Non-RSS webpages (scraped automatically, slower): https://example.com/news
2. Plugin Settings UI
You can adjust these values in the LM Studio plugin settings panel:
| Setting | Default | Description |
|---|---|---|
retentionDays | 90 | How long to keep articles (30, 60, 90, or 365 days). |
refreshIntervalHours | 6 | How often the plugin automatically fetches new RSS items in the background. |
embeddingModel | nomic-embed-text-v1.5 | The exact model ID loaded in your LM Studio server. |
topK | 8 | How many unique articles to return per search. |
chunkSize | 800 | Length (in characters) of each text chunk used for embedding and search. |
embedBatchSize | 16 | How many chunks are sent to the embedding model per request (higher = faster but more VRAM). |
maxConcurrentFeeds | 8 | How many RSS/Atom feeds are fetched at the same time. |
htmlLinkPatterns | news,article,post,blog,stories,press | Comma-separated URL keywords used when scraping non-RSS pages. |
scrapeFullText | Off | For HTML-scraped pages, also download each article's full body text (slower refresh, richer content). |
retentionBasis | publish_date | Whether old entries are removed by publish date or cache date. |
refreshOnStartup | On | Automatically fetch feeds when the plugin loads. |
🛠️ LM Studio Tools
The plugin exposes these tools to the LLM:
refresh_rss_feeds — Force-refreshes all feeds listed in feeds.txt, indexes new items, generates embeddings, and applies the retention policy. Returns a report of successes, failures, and added articles.search_rss_cache — Hybrid (keyword + semantic) search, optionally filtered by days, feed (URL substring), or author.list_rss_feeds — Lists all configured feed URLs plus cache stats: entry/chunk counts, missing embeddings, DB size, and last refresh report.get_recent_cache_articles — Returns the most recent cached articles. Use this to see what news is available without guessing keywords.add_rss_feed — Adds an RSS/Atom feed URL or webpage URL to feeds.txt.remove_rss_feed — Removes a feed URL from feeds.txt (cached articles kept until pruned).test_feed — Checks whether a URL is a working RSS/Atom feed or scrapeable HTML page, and how many items it exposes.prune_removed_feeds — Deletes cached articles whose source feed is no longer in feeds.txt.reindex_missing_embeddings — Re-embeds chunks stored as text-only while the embedding model was offline.🧠 How It Works
fast-xml-parser (with a legacy regex fallback). If a URL returns HTML, it uses Cheerio to scrape the page for article links.maxConcurrentFeeds). HTML pages are fetched one-by-one with a 1.5–3 second delay to avoid anti-bot firewalls. Requests retry with backoff on 429/403/5xx./v1/embeddings endpoint in batches. If the model is offline, chunks are stored text-only and reindex_missing_embeddings completes them later.cache.db). Newer items list first; retention purges by publish or cache date.🐛 Troubleshooting
Search returns 0 matches
Ensure your embedding model is loaded and the server is running. If the model is offline, the plugin falls back to keyword search — but if your articles don't contain the exact keywords, they won't be found. Run list_rss_feeds to check that chunks > 0, or reindex_missing_embeddings once the model is loaded.
HTTP 403 Forbidden on a feed Some websites block automated requests. The plugin uses a Chrome User-Agent, throttles HTML requests, and retries with backoff, but strict firewalls might still block it. Try an alternative RSS URL or an RSS Bridge service.
HTTP 429 Too Many Requests
Sites like Reddit aggressively block Node.js fetch requests. If this happens, use an RSS Bridge URL (e.g., https://rss-bridge.org/bridge01/?action=display&bridge=RedditBridge&format=Atom) instead of the native .rss link.