LATEST / OLLAMA
29 SEP 2026 · DECISION MODELS ARRIVE31 AUG 2026 · CLOUD BILLING MOVES TO TOKEN RATES09 JUL 2026 · $65M SERIES B ANNOUNCED
Company / AI + Developer toolsThe friction issue / 01

Ollama made AI easier to run. Now it sells the room to grow.

The founders who helped tame Docker turned their attention to open AI models. Ollama’s wager: make the first experiment easy, then keep the same tools when the laptop runs out of room.

A downloadable AI model is a peculiar sort of gift. You have been handed the machinery, but someone has misplaced the instructions for the workshop. The files are there. The questions are there. Between them sit hardware choices, memory limits and software that would quite like to know which version of everything else you installed.

Ollama built a company in that gap. Its proposition is pleasingly modest: make open models easier to run. Download the software, choose a supported model, and get to the first useful exchange. For a developer, that can mean a command such as ollama run llama3.2. A research artifact begins to behave like an everyday tool.

The useful bits / 30 seconds
  • What it does: runs and manages open AI models locally, with a hosted cloud option for larger workloads.
  • Who it serves: developers, researchers and teams building with AI, plus people who want a desktop chat interface.
  • What it charges: no Ollama inference fee for local models; paid cloud credits and subscriptions.
  • The catch: convenience cannot manufacture memory, guarantee a correct answer or erase a model’s license.

The same itch, ten years later

Jeffrey Morgan and Michael Chiang had encountered this kind of problem before. They met in college and built Kitematic, a tool that made Docker easier to run. Docker bought it in 2015. Their work became Docker Desktop, launched in 2016. The pattern matters more than the résumé: useful technology existed, but the distance between knowing about it and using it was unnecessarily long.

Ollama applied that instinct to open models in 2023. The company’s lineage gives its product choices a certain logic. A developer trying a model should not have to become an inference specialist before discovering whether it helps with the task. Every preliminary chore is another opportunity to abandon the experiment.

Ollama co-founders Jeffrey Morgan and Michael Chiang
Same founders, another installation headache. Jeffrey Morgan and Michael Chiang built Kitematic before Ollama. Photo: David Paul Morris / Ollama.

Its distribution is already substantial. Ollama reported 8.9 million monthly active developers in July 2026, alongside use inside 85% of the Fortune 500. Read the second number carefully. An employee trying a tool, a team adopting it and a corporation buying a production service are different events. The figure describes reach; it does not establish 425 paying enterprise customers.

The founders also announced $88 million raised in total. A $65 million Series B was led by Theory Ventures. That capital gives a company famous for running models on someone else’s computer resources to build a business running them on its own cloud.

“Open models should be easy to run”Jeffrey Morgan, co-founder and CEO, July 2026

The model is the guest. Ollama manages the house.

Ollama occupies the practical layer between model publishers and the people building applications. It supplies a runtime, model downloads and management, developer interfaces and a desktop experience. The intelligence comes from the selected model. Choosing Gemma, Qwen or Llama can change the answer even when the surrounding Ollama workflow remains familiar.

The distinction explains both its usefulness and its limits. A model can draft code, summarize notes or answer questions. Ollama helps make that model available to a program or a person. It does not automatically turn a weak response into a sound one. Nor does the runtime’s MIT license grant every model identical commercial permissions.

The product has widened beyond the terminal. In July 2025, Ollama introduced a chat interface for macOS and Windows, with model downloads and file drag and drop. A user could bring text or PDFs into a conversation, or send images to a model that supports image inputs. That makes the first experiment accessible to people who would rather not spend their afternoon debating a shell command.

Official Ollama desktop chat application screenshot
The terminal has acquired a sitting room. Ollama’s desktop app puts model chat and file attachments behind a familiar window. Image: Ollama.

For builders, there are Python and JavaScript clients and an API. Embedding models provide vectors for search and retrieval applications. Structured outputs can constrain supported responses to a JSON schema: handy when extracting fields from documents, where a charming paragraph is less useful than a record your software can parse.

Customization also needs a precise description. A Modelfile can set a system prompt and model parameters, or incorporate supported adapters. Giving an assistant a house style is different from training new knowledge into its weights. Ollama makes configuration approachable; it does not make the distinction disappear.

The laptop eventually asks for help

Local inference has an attractive economic property: experimenting does not add an Ollama charge for every prompt. Once a model is downloaded, supported local workflows can also operate without an ongoing cloud connection. The costs have moved to equipment, power and time. A slow machine makes that last item conspicuous.

The physical boundary arrives quickly with larger models, longer documents or multiple simultaneous requests. Ollama’s documentation explains that context length and parallel processing increase memory requirements. Local compute is a finite budget, even when the software invoice reads zero.

Cloud models entered preview in September 2025. Their appeal was continuity: use familiar tools while the model runs on data-center hardware. Local and hosted options sit in the same product, but the data boundary changes. A cloud request leaves the computer. Ollama says it does not log or train on cloud prompts and responses; that promise differs from keeping inference entirely local.

The alternatives sit at different points in the stack. LM Studio offers another desktop route. Using llama.cpp directly gives technical users a closer relationship with the engine; Ollama itself uses llama.cpp technology. vLLM addresses model serving. Hosted APIs offer managed access without requiring local hardware. The sensible comparison starts with the workload, rather than treating all four as interchangeable boxes.

The bill that changed their minds

In August 2026, Ollama made a revealing change. Customers had found GPU-time billing difficult to predict, the company said. New cloud plans moved to per-token rates with included monthly credits. The lesson is prosaic and useful: a measurement that describes infrastructure elegantly can still make a customer’s budget difficult to explain.

At the October 1 check, Pro costs $20 a month and includes $60 of usage credits. Max costs $100 with $300 included. Team costs $500 with $1,000 shared across unlimited users. These are usage allowances, not cash rebates. Model rates determine how far they go, unused included credits do not roll over, and additional usage costs money.

Cloud subscriptions / monthly / October 2026
Pro$20$60 usage
Max$100$300 usage
Team$500$1,000 shared

Local inference remains free of Ollama usage fees. Enterprise pricing is custom.

The business model follows the product’s boundary. Offer local experimentation freely, then charge for hosted capacity and organizational conveniences. Teams may value pooled billing and administration as much as another model. Free accounts can also buy cloud credits without a subscription, leaving a smaller commitment available to an occasional user.

What happens underneath the easy command

The engineering is less dainty than the interface. In September 2025, Ollama described a scheduling change that measured model memory requirements more exactly. Earlier estimates could over-allocate memory and produce crashes. Improving that machinery also helped place more work on GPUs and distribute it across multiple devices.

One published test makes the point concrete. For Gemma 3 12B at 128k context on an NVIDIA RTX 4090, Ollama reported generation rising from 52.02 to 85.54 tokens per second. That is a result for one configuration, not a promise about your laptop. The interesting detail is that fitting the work onto the right hardware changed performance without requiring a different task.

One reported test / tokens per second
Before
52.02
After
85.54
The GPU had more to give. Ollama’s September 2025 test: Gemma 3 12B, 128k context, one RTX 4090. Configuration-specific, company-reported.

Its NVIDIA DGX Spark partnership addresses the same practical territory: chat, document processing, code and multimodal workloads. Partnerships with model publishers including IBM and Meta help bring their releases into Ollama’s library. The expertise lies in packaging, compatibility and inference engineering - work that matters most when it becomes difficult to notice.

Copy the experiment, not the slogan

A useful way to try Ollama is to choose one bounded job. Extract a few fields from sample documents. Draft a function with tests. Summarize a folder of non-sensitive notes. Start with a small supported model and keep reference answers so you can distinguish plausible output from correct output.

Then measure the whole task: latency, accuracy, retries and human correction. A cheap response becomes expensive if someone must repair it. Increase context or model size only when the evaluation earns it. If local memory becomes the obstacle, compare a hosted run before buying hardware.

That approach suits experimentation, offline work and applications whose chosen models meet the task. It becomes less comfortable when the workload needs more memory or throughput than you have, a deployment requires controls you have not implemented, or errors are too costly for the available model. An easy installation cannot settle those decisions.

Ollama’s repeatable idea is to make the first useful attempt less troublesome. Its founders have built around that idea twice. The open model may draw the attention; the quiet reduction in chores is what gives someone a reason to try it before lunch.