In Chapter 3.2 we talked about the prototype that arrives in five minutes, and I ended it with a promise: next chapter, we work without touching the cloud at all. Here it is.
The case for a local LLM usually gets compressed into two sentences: “my code never leaves this machine” and “I stop paying per token”. Both are true. The third sentence nobody says out loud is this: the moment you unplug from the grid, the power, the machine and the maintenance are yours.
Using a hosted model is like buying electricity from the grid. You plug in, it works, at the end of the month you pay what the meter says. Running a model locally is putting a generator in the garden: you buy the machine up front, you fill the tank, and how many sockets it can feed is decided by the machine, not by you. Nobody audits you. But when the lights go out, there is no utility company to blame either.
So let us go through the bill line by line: which model actually fits, why your context window gets trimmed without telling you, what the licence really says, and what all of this means on a server running PHP.
What the Generator Fixes, and What It Does Not
Let us align expectations first, because half of the local-model argument comes from the wrong ones.
What it fixes: client code, database schemas and customer data never leave the machine. That is the technical counterpart of the confidentiality clause you signed, and it means you no longer have to take “we do not train on your data” on faith. It works with the network down. There is no monthly quota, so you stop weighing whether a prompt is worth re-running.
What it does not fix: latency. In the cloud, the provider’s data centre decides how fast you get an answer; at home, your card does. You also do not get the best model in its class; the gap between downloadable weights and closed models is narrowing, but it is not zero. And the sneakiest one: context is no longer free. Sending 200k tokens to a hosted model costs money. Doing it at home costs room.
A generator will light the house. Whether it also runs the air conditioning is a question about its rating.
Line One: How Big Is the Generator?
Picking a model is not a matter of taste, it is a fitting problem. Download sizes below come from the Ollama library; parameter and context figures come from the official model cards (as of 27 September 2026).
| Model | Total / active parameters | Native context | Licence | Ollama download |
|---|---|---|---|---|
| qwen3-coder:30b | 30.5B / 3.3B | 262,144 tokens | Apache-2.0 | 19 GB |
| qwen3-coder:480b | 480B / 35B | 262,144 tokens | Apache-2.0 | 290 GB |
| llama4:16x17b (Scout) | 109B / 17B | 10M (announced) | Llama Community License | 67 GB |
| llama4:128x17b (Maverick) | 400B / 17B | not announced | Llama Community License | 245 GB |
The message is not in the columns, it is in the gap between them. The 30B build of Qwen3-Coder is a mixture-of-experts model: 30.5 billion parameters in total, but only 3.3 billion active per token. You load all of it into memory and a fraction of it does the arithmetic. A 19 GB download looks reasonable on its own. The catch is that the context window sits on top of those 19 GB.
The hard part: download size is not running size. The “19 GB on disk, 24 GB on the card, we are fine” arithmetic breaks the moment you widen the window.
Line Two: How Many Sockets Does It Feed?
Here is the most expensive sentence in this chapter. Ollama picks your context window for you, based on available VRAM. Its own documentation lists these defaults:
- under 24 GB of VRAM: 4k tokens
- 24–48 GB: 32k tokens
- 48 GB and above: 256k tokens
Now put that next to the table above, where Qwen3-Coder’s native context reads 262,144 tokens. On a 16 GB card you will be running that model through a 4,096-token window. Same model. Your window.
The irritating part is how quietly it happens. No error, no warning. The agent simply forgets what was in the first file by the time it opens the third, and you conclude that local models are not ready. It looks a lot like the problem I described in the goldfish-memory article (in Turkish), except the cause is not the model. It is your setting.
Ollama’s documentation is direct about it: tasks that need large context, such as search, agents and coding tools, should be set to at least 64,000 tokens. Start the server accordingly:
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
If you would rather set it per request, the field is options.num_ctx. The same body also carries keep_alive, which defaults to five minutes: after that the model is evicted from memory and reloaded on your next call. If you have ever wondered why the first question of the morning takes so long, that is usually the answer.
curl http://localhost:11434/api/chat -d '{
"model": "qwen3-coder:30b",
"messages": [{"role": "user", "content": "Why is this SQL slow?"}],
"stream": false,
"keep_alive": "30m",
"options": { "num_ctx": 64000, "temperature": 0.2 }
}'
The point: with a local model the dial that matters is not the model name. It is the width of the window and how long the weights stay resident.
Line Three: Who Actually Reads the Licence?
“Open model” and “open source model” are not the same thing, and the difference shows up in the project you hand to a client.
Qwen3-Coder ships under Apache-2.0, as stated on its Hugging Face model card. Familiar, predictable, nothing for a legal team to argue about. The model comes from the Qwen team at Alibaba; the 480B build was announced on 22 July 2025.
Llama is a different situation. Meta distributes the weights openly, but under the Llama Community License, which is neither Apache nor MIT. In a post dated 18 February 2025, the Open Source Initiative states that this licence does not meet the Open Source Definition: it restricts the freedom to use the model for any purpose, discriminates between users, and limits fields of endeavour.
None of that matters while you are experimenting on your own machine. It matters on the day you install the model on a client’s server and send an invoice. Open the licence file before you sign the contract; it belongs in the same category as checking end-of-support dates before an upgrade.
The Engine Is Not the Client
In Chapter 3.1 (in Turkish) we established that the editor is not the only door. On the local side the split is sharper still: the layer that runs the model and the layer that talks to it are separate pieces of software, chosen separately.
The inference engine
llama.cpp sits underneath most of this: an MIT-licensed inference engine written in C/C++, supporting quantization from 1.5-bit up to 8-bit. The project was started by Georgi Gerganov (@ggerganov), and on 20 February 2026 the ggml team joined Hugging Face. The announcement says the project “will continue to be 100% open-source and community driven”, with the team keeping “full autonomy and leadership on the technical directions”. When there is an organisation paying salaries behind a piece of infrastructure, that is good news for anyone running it in production.
Ollama (MIT) puts a package manager and an HTTP server on top of that engine; the current release is 0.34.4, published on 23 September 2026. LM Studio gives you the most comfortable desktop interface of the three, but it is closed source: its terms grant free use for personal and internal business purposes and forbid redistributing the software.
The client side
Cline (Apache-2.0, Cline Bot Inc.) runs in VS Code and JetBrains and supports Ollama and LM Studio as providers directly, with http://localhost:11434 as the base URL. Its documentation suggests turning on “Use Compact Prompt” and keeping tasks small for local use. They know about the context problem too.
Aider (Apache-2.0) is Paul Gauthier‘s terminal pair-programming tool, and it connects to almost any model, local ones included. The latest release on PyPI is 0.86.2, dated 12 February 2026.
There used to be a third name here: Continue. Same category, Apache-2.0, Continue Dev, Inc. behind it. The repository today opens with this: “no longer actively maintained and is read-only for all users”. The lesson is not about that tool. Do not build your workflow on a single extension. Keep the engine standard and the client replaceable.
The hosted agents — OpenAI Codex, Anthropic Claude Code, Google Antigravity — are not outside this table so much as on the other side of it. Three different ecosystems do the same job, and the choice between them is practical, not ideological.
| Layer | Example | Licence | Runs on |
|---|---|---|---|
| Inference engine | llama.cpp | MIT | Your hardware |
| Engine + package manager | Ollama 0.34.4 | MIT | Your hardware |
| Desktop interface | LM Studio | Closed source, free for personal and internal business use | Your hardware |
| Editor agent | Cline | Apache-2.0 | Local or hosted model |
| Terminal agent | Aider 0.86.2 | Apache-2.0 | Local or hosted model |
| Hosted agent | Codex, Claude Code, Antigravity | Proprietary service | The provider’s infrastructure |
What This Is Good For on the PHP Side
The most honest use of a local model in a PHP project is not “let it write my code”. It is processing text that must not leave the building: order notes, support tickets, error logs, contract text. When you do not want that material going to a hosted API, this is exactly what the generator is for.
The pattern is the same one from the Crawl4AI write-up: PHP talks to a local service over HTTP and the heavy work happens there. The class below is framework-agnostic, plain PHP 8.3. Drop it into a DI container if you have one, or just new it; it behaves the same either way. It has been through php -l.
<?php
declare(strict_types=1);
final class OllamaClient
{
public function __construct(
private readonly string $baseUri = 'http://127.0.0.1:11434',
private readonly string $model = 'qwen3-coder:30b',
private readonly int $numCtx = 64000,
private readonly int $timeout = 180
) {
}
/**
* @param array<int, array{role: string, content: string}> $messages
*/
public function chat(array $messages): string
{
$payload = [
'model' => $this->model,
'messages' => $messages,
'stream' => false,
'keep_alive' => '30m',
'options' => [
'num_ctx' => $this->numCtx,
'temperature' => 0.2,
],
];
$ch = curl_init($this->baseUri . '/api/chat');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_TIMEOUT => $this->timeout,
CURLOPT_HTTPHEADER => ['Content-Type: application/json'],
CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR | JSON_UNESCAPED_UNICODE),
]);
$raw = curl_exec($ch);
$code = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$err = curl_error($ch);
curl_close($ch);
if ($raw === false) {
throw new RuntimeException('Ollama unreachable: ' . $err);
}
if ($code !== 200) {
throw new RuntimeException('Ollama HTTP ' . $code . ': ' . substr((string) $raw, 0, 300));
}
$data = json_decode((string) $raw, true, 512, JSON_THROW_ON_ERROR);
if (!isset($data['message']['content'])) {
throw new RuntimeException('Unexpected response body.');
}
return trim((string) $data['message']['content']);
}
}
Three details matter. num_ctx is set explicitly, because trusting the server default is the trap described above. keep_alive is raised to half an hour, because in a background queue a model that reloads on every job doubles your wall clock. And the timeout is 180 seconds: do not expect the three-second replies you are used to from a hosted API.
Ollama also exposes an OpenAI-compatible endpoint at http://localhost:11434/v1/chat/completions. If you already have an OpenAI client wired up, changing the base URL is often the whole migration.
Generator or Grid?
| Situation | Choice | Why |
|---|---|---|
| Client code or personal data is involved | Local | Nothing leaves the machine; the contract clause is met technically |
| Repetitive, patterned work (log triage, text cleanup) | Local | High volume, simple task, no quota |
| Designing architecture, large refactors | Hosted | Needs wide context and the strongest reasoning available |
| Under 16 GB of VRAM | Hosted | A 4,000-token window is not enough for agent work |
| Air-gapped or offline environment | Local | The only option that runs |
These two are not alternatives. Serious setups run both: sensitive data locally, heavy thinking in the cloud. Keeping that switch in your own hands means writing the client independently of the provider, which is precisely why the base URL in the class above is configurable.
There Is No Free Electricity
A local model does not remove the bill, it changes the line items. You pay in VRAM instead of tokens, in context instead of quota, in your own maintenance instead of the provider’s downtime. What you get back is real and not small: you know where your code goes.
Three things to write down before you install anything. Set the window by hand. Read the licence before you download the weights. Keep the engine separate from the client. Do all three and the generator lights the house. Skip them and you are left with a 19 GB file and an agent that forgets what was in the third file.
Next chapter (4.1) turns the wheel: mega-prompting, or making the model the lead architect of your project. Wherever it runs, what you tell it is what decides the outcome. The full series index lives in the Vibe Coder’s Handbook (in Turkish).
Stay with the technology, and read the bill before you sign it.