Can connect to several existing backends like Anthropic, Cohere, OpenAI, NvidiaNIM, MistralAI etc,
and host models on its own - see the Cortex section on the screenshot below - showing Jan downloaded and hosting locally Llama3 8b q4 and Phi3 medium (q4).
Pros (What I liked):
Intuitive interface
Ability to experiment with model temperature, topp, frequency and presense penalties and system prompts.
Provides API server
Cons:
Somehow slow on my ubuntu-based os. On windows it did run ok.
Can connect to many backends, but all of them are managed. Would be nice to use Ollama option.
Not many variants of the models available for self-hosting in Cortex. Not too many quantizations options either.
Yes, Huggingface gguf is awesome. But I wanted
to reuse what ollama already downloaded loaded into VRAM
Vane is one of the more pragmatic entries in the “AI search with citations” space: a self-hosted answering engine that mixes live web retrieval with local or cloud LLMs, while keeping the whole stack under your control.
Quick overview of most prominent UIs for Ollama in 2025
Locally hosted Ollama allows to run large language models on your own machine, but using it via command-line isn’t user-friendly.
Here are several open-source projects provide ChatGPT-style interfaces that connect to a local Ollama.
That’s very exciting!
Instead of calling copilot or perplexity.ai and telling all the world what you are after,
you can now host similar service on your own PC or laptop!
Subscribe
Get new posts on AI systems, Infrastructure, and AI engineering.