AI - Page 3 - Rost Glukhov | Personal site and technical blog

Ollama’s GPT-OSS models have recurring issues handling structured output, especially when used with frameworks like LangChain, OpenAI SDK, vllm, and others.

Constraining LLMs with Structured Output: Ollama, Qwen3 & Python or Go

Large Language Models (LLMs) are powerful, but in production we rarely want free-form paragraphs. Instead, we want predictable data: attributes, facts, or structured objects you can feed into an app. That’s LLM Structured Output.

Memory allocation and model scheduling in Ollama new version - v0.12.1

Here I am comparing how much VRAM new version of Ollama allocating for the model vs previous Ollama version. The new version is worse.

Ollama Enshittification - the Early Signs

Ollama has quickly become one of the most popular tools for running LLMs locally. Its simple CLI, and streamlined model management have made it a go-to option for developers who want to work with AI models outside the cloud. But as with many promising platforms, there are already signs of Enshittification:

Locally hosted Ollama allows to run large language models on your own machine, but using it via command-line isn’t user-friendly. Here are several open-source projects provide ChatGPT-style interfaces that connect to a local Ollama.

Popularity of Programming Languages and Software Developer Tools

The Pragmatic Engineer letter published a couple days ago survey statistics of the popularity of programming languages, IDEs, AI tools and other data for mid-2025.

NVIDIA DGX Spark - new little AI supercomputer

Nvidia is about to release NVIDIA DGX Spark - little AI supercomputer on blackwell architecture with 128+GB unified RAM and 1 PFLOPS AI performance. Nice device to run LLMs.

Reranking documents with Ollama and Qwen3 Reranker model - in Go

Since standard Ollama doesn’t have a direct rerank API, you’ll need to implement reranking using Qwen3 Reranker in GO by generating embeddings for query-document pairs and scoring them.

Comparison of Hugo Page Translation quality - LLMs on Ollama

In this test I’m comparing how different LLMs hosted on Ollama translate Hugo page in English to German. Three pages I tested were on different topics, had some nice markdown with some structure: headers, lists, tables, links, etc.

Reranking texts with Ollama and Qwen3 Embedding LLM - in Go

This little Reranking Go code example is calling Ollama to generate embeddings for the query and for eache candidate document, then sorting descending by cosine similarity.

LLM Performance and PCIe Lanes: Key Considerations

How PCIe Lanes Affect LLM Performance? Depending on the task. For training and multi-gpu inferrence - perdormance drop is significant.

Convert HTML content to Markdown using LLM and Ollama

In the Ollama models library there are models that able convert HTML content to Markdown, which is useful for content conversion tasks.

Search is best for quick, straightforward information retrieval using keywords.
Deep Search excels at understanding context and intent, delivering more relevant and comprehensive results for complex queries.

Will list here some AI-assisted coding tools and AI Coding Assistants and their nice sides.

Using LLMs is not very expensive, might be no need to buy new awesome GPU. Here is a list if LLM providers in the cloud with LLMs they host.

Test: How Ollama is using Intel CPU Performance and Efficient Cores

I’ve got a theory to test - if utilising ALL cores on Intel CPU would raise the speed of LLMs? This is bugging me that new gemma3 27 bit model (gemma3:27b, 17GB on ollama) is not fitting into 16GB VRAM of my GPU, and partially running on CPU.

AI

Ollama GPT-OSS Structured Output Issues

Constraining LLMs with Structured Output: Ollama, Qwen3 & Python or Go

Memory allocation and model scheduling in Ollama new version - v0.12.1

Ollama Enshittification - the Early Signs

Chat UIs for Local Ollama Instances

Popularity of Programming Languages and Software Developer Tools

NVIDIA DGX Spark - new little AI supercomputer

Reranking documents with Ollama and Qwen3 Reranker model - in Go

Comparison of Hugo Page Translation quality - LLMs on Ollama

Reranking texts with Ollama and Qwen3 Embedding LLM - in Go

LLM Performance and PCIe Lanes: Key Considerations

Convert HTML content to Markdown using LLM and Ollama

Search vs Deepsearch vs Deep Research

AI Coding Assistants comparison

Cloud LLM Providers

Test: How Ollama is using Intel CPU Performance and Efficient Cores