README
Text-only Documentation
Looking for the illustrated user documentation? https://github.com/ceveyne/find-image-docs
Looking for the runtime engine? See https://github.com/ceveyne/qwen3-vl-embedding/releases
LM Studio Plugin — Next-level image search and information retrieval based on multimodal embeddings
find-image supports text-based, image-based, and metadata-based queries using Qwen3-VL-Embedding.
find-image is an advanced LM Studio plugin for searching local image stores and generation history. It helps finding images, prompts, source files, and project entries even if you remember only part of what you are looking for. 🧠
The plugin combines metadata search with optional multimodal embedding search. Exact text matches still matter, but visual embeddings allow results to be found by image content instead of only by prompt or filename text.
Images can be tagged and found with exact tag filters.
When used together with draw-things-chat or user-docs, search results can be shown with visual previews and can be referenced in follow-up analysis, detection, generation, edit, image-to-image, image-to-video, or post-processing requests.
The local index can include:
The multimodal index accepts PNG, JPG/JPEG, TGA, BMP, PSD, GIF, HDR, PIC, PPM, and PGM images. WebP, TIFF/TIF, HEIC, and HEIF are not supported.
Setup is an easy 3-step process:
Your model choice affects search quality, VRAM usage, and how long indexing takes. The smallest, fastest variant (2B.Q5_K_S) already gives useful results. If you have a powerful computer and want results close to the unquantized original, choose 8B.Q8_0.
💡 Each model variant creates its own embeddings. Want to try a few models? Start with a dedicated test folder of images. That keeps indexing fast and your experiments easy to compare.
⚠️ Do not put Qwen3-VL-Embedding models in
~/.lmstudio/modelsor another LM Studio model directory. LM Studio does not support them yet, but it does not simply ignore them either. They can interfere with supported Qwen3-VL models such asqwen/qwen3-vl-4b,qwen/qwen3-vl-8b, orqwen/qwen3-vl-30b.
👍 The default Multimodal Model Path is ~/.lmstudio/extensions/models/Qwen3-VL-Embedding-8B-GGUF. Use that location and there is nothing else to configure in the plugin.
💡 If your model directory contains more than one model, include the file name too, for example: ~/.lmstudio/extensions/models/Qwen3-VL-Embedding-8B-GGUF/Qwen3-VL-Embedding-8B.Q8_0.gguf. If it contains only one model, the directory path is enough.
There are plenty of models to choose from. These instructions were tested with the 8B models from mradermacher (https://huggingface.co/mradermacher/Qwen3-VL-Embedding-8B-GGUF) and the 2B models from DevQuasar (https://huggingface.co/DevQuasar/Qwen.Qwen3-VL-Embedding-2B-GGUF).
💡 Whichever quantization you choose, always use an f16 mmproj, for example Qwen3-VL-Embedding-8B.mmproj-f16.gguf.
Upstream llama.cpp does not fully support Qwen3-VL-Embedding yet, so LM Studio's built-in llama-server cannot run these models. find-image uses its own compatible llama-server instead. Download the macOS binaries from the releases page, or build them yourself.
👍 The default GGUF Embedding Binary Path is ~/.lmstudio/extensions/backends/qwen3-vl-embedding/llama-server. Place it there and the plugin needs no further configuration.
💡 LM Studio also stores its own llama-server binaries in ~/.lmstudio/extensions/backends.
⚠️ macOS quarantines downloaded files by default. To let it use the downloaded, ad hoc-signed llama-server binaries, remove that quarantine attribute, just as you would for Draw Things gRPCServerCLI-macOS:
xattr -dr com.apple.quarantine ~/.lmstudio/extensions/backends/qwen3-vl-embedding/ && xattr -l ~/.lmstudio/extensions/backends/qwen3-vl-embedding/libllama-server-impl.dylib
Change the path if you stored the binaries somewhere else.
🚧 You can skip this step if you build the binaries yourself.
To install the plugin, select Run in LM Studio.
⚠️ find-image also works on its own, but it is much more useful when you can see results in chat and keep working with the images you find. For that, use the current version of draw-things-chat or user-docs. Update them before you start using find-image.
The LM Studio Hub does not yet offer an easy update process. It will not notify you about new versions, and you cannot automatically replace an installed plugin.
🦺 Delete the old plugin version manually, restart LM Studio, then install the update.
The good news is that your plugin settings and valuable embeddings stay intact.
💡 Tip: Watch the documentation project on GitHub to get notified about updates. 👀
~/.find-image
~/.find-image/data/generation_index_cache.json ~/.find-image/data/multimodal_embeddings.sqlite3
The plugin settings are mostly self-explanatory and are covered above in the setup instructions. Notes:
find_image call, or not yet indexed at all, are embedded.Set this to 0 to pause new embeddings or 100 to index the entire backlog. With a limited value, each run rotates through enabled image directories, chat sources, and Draw Things projects so a large folder does not delay every other source.
For find_image to work well, everything you want to find must be indexed first. Text information such as filenames and generation metadata takes only a few milliseconds to index, but image data can take several seconds. On a slow machine with a large model such as Qwen3-VL-Embedding-8B.Q8_0 and very large images, it can take up to a minute per image. The recommendation is therefore: start small.
Qwen3-VL-Embedding-8B.Q8_0 to see whether it better meets your requirements.After testing and choosing the model that works best for you, set Max New Multimodal Embeddings per Run to 100 (unlimited) to index all desired images and project files. This can take time, but it can run in the background as long as foreground work is not also heavily using VRAM and the GPU.
💡 General note about plugin behavior: The qwen3-vl-embedding llama-server starts for each tool call and stops when that call is complete. It does not run permanently in the background, so it releases all required resources when it has finished.
⚠️ Expert tip for researchers: Each model creates a separate index. If you expect to experiment with different models (quantizations), specify the exact GGUF file in the Multimodal Model Path right from the start. This creates a clear, isolated setup. Do not specify only a directory in Multimodal Model Path and then try different GGUF models with the same path. Doing so produces a mixed index with unpredictable search results.
The find_image tool interface has four fields:
aN, vN, iN, pN) or absolute image path; use it for "like this image", style references, or changes to a shown imagetarget, positively describe additions or changes; without target, write a complete description. Model:, LoRAs:, Size:, Source:, Origin:, and Timestamp: are hard AND filters.Use one, two, three, or all four fields in each search, depending on what you want to find or achieve.
| Goal | Recommended Parameters | Why? |
|---|---|---|
| "Find visually similar images" | {target} | The reference image ranks the complete corpus by visual similarity. |
| "Same style, but different subject or scene" | target + query | Keep the reference as the visual anchor; positively describe the change, for example people in front of a brick wall. |
| "More from this creative direction" | target + includeMetadata | Image pixels and target metadata jointly rank the complete corpus by visual and generation-context similarity. |
| "Reference style plus a specific new concept" | target + query + includeMetadata | Combines visual reference, requested changes, and target metadata as ranking signals. |
| "Thematic search without a reference image" | {query} | Uses a complete descriptive text query across the full collection. |
| "Require specific generation metadata" | query with Model:, LoRAs:, Size:, Source:, Origin:, or Timestamp: | These hard AND filters limit candidates before retrieval and reranking. |
| "Prompt or metadata similarity without visual similarity" | target + excludeImage + includeMetadata | Uses reference metadata as a text-based similarity signal while deliberately omitting image pixels. |
The tag_image tool allows you to manage persistent tags for indexed images:
find_image as a distinct filter to help you organize your images.This is the part that makes find-image different from a plain filename or prompt search. It also explains why the same search can behave differently depending on whether you provide an image, text, or both.
When find-image indexes an image, it looks at the image itself and, when available, its generation metadata. In other words, it considers both what the image looks like and how it was made. It turns this information into an embedding: a long list of numbers that represents the image in the search index.
Think of an embedding as a pin on a very large map. Images and descriptions that are alike end up close together. When you search, find-image creates a pin for your query and looks for the closest pins in your indexed collection. The technical name for this comparison is cosine similarity; the result becomes the match score you see.
For example, attach an image with metadata and search with the image plus its metadata only:
find_image: { "target": "a1", "query": "", "includeMetadata": true }
find-image then looks for images that are close in both appearance and generation context. This is useful when you want more images from the same visual style, session, or creative direction.
If your attached image is already in the index, find-image leaves it out of the results. Otherwise, the first result would often just be the same image again. By default, it also filters out any byte-identical copies for the same reason.
So in this mode, similar means similar in both the pixels and the metadata, not merely one or the other.
You can then steer the search by adding your own description. For example, attach a portrait and add cinematic neon lighting. The attached image keeps the visual search grounded, while your text guides it toward a particular idea or mood.
That is the useful bit about multimodal embeddings: an image, a plain-language description, and generation metadata can all point toward the same area of the index. You can search with a reference image, describe what you want in words, or use the metadata from an AI-generated image. Mix those signals when you want more control over the results.
target and includeMetadata rank results from the complete corpus by visual and generation-context similarity.query, such as Model: and Size:, are hard AND filters that restrict candidates before retrieval.excludeImage only when prompt or metadata similarity should matter without target image pixels.target, describe desired additions or replacements positively instead of reconstructing the reference image in text.target, use a complete description for thematic text search.target + query + includeMetadata combines visual reference, requested changes, and metadata similarity without restricting results to identical metadata.Score ≠ Quality. Score = Embedding Coverage.
A lower score does not necessarily mean worse results. Signals change what ranks first, while structured query metadata determines which candidates are eligible.
| Feature | find-image | index-image |
|---|---|---|
| Text-based queries | ✅ | ✅ |
| Image-based queries | ✅ | ⛔️ |
| Multimodal queries | ✅ | ⛔️ |
| Search by metadata | ✅ | ✅ |
| Semantic search | ✅ | ✅ |
| Multilingual search | ✅ | (✅) depends on model |
| Setup | 3-step | 2-step |
| Python venv | none | none |
| Model supported by LM Studio | ⛔️ | ✅ |
| llama-server | ✅ | ✅ |
| Query performance | 🚀 | 🚀 |
| Embedding performance | seconds | milliseconds |
Conclusion: If you need neither image-based nor multimodal search, and do not require strong multilingual search, index-image is the better choice. In all other cases, find-image is recommended.
See CHANGELOG.md for version history and release notes.
MIT