
Adapting Open-Weight Models: When Retrieval, Fine-Tuning, or Distillation Wins
A team has an open-weight model in production. The outputs are wrong in a specific, repeatable way. Someone says the words "fine-tuning run," and suddenly…
Read reportMethods for selecting, assembling, retaining, and supplying evidence or state to AI systems, including retrieval-augmented generation and long-context designs.
Tagged articles
19 articles in this tag.

A team has an open-weight model in production. The outputs are wrong in a specific, repeatable way. Someone says the words "fine-tuning run," and suddenly…
Read report
The agent that answered cleanly for ten minutes starts contradicting a decision it made at turn twelve. It asks again for a fact you supplied in the…
Read report
Your pager fires at 2:14 a.m. A customer has posted a screenshot: your assistant told someone to adjust a medication dose. You open the trace, find the…
Read report
The hard part was never the model call. It is everything you have to build around it.
Read report
The cluster is provisioned, the dashboards are green, and the inference bill is still climbing while latency drifts. That combination — healthy…
Read report
The prompt was precise. The model was capable. The change still came back wrong.
Read report
A page can rank, read beautifully, and still lose the answer. The reason is structural: many retrieval systems never see your page whole. They see…
Read report
The retrieval worked. The document was in the prompt. The model still answered wrong.
Read report
That gap — between what retrieval found and what the model used — is where context packing lives. It is the assembly stage between retrieval and…
Read report
The answer is wrong, so someone rewrites the prompt. Still wrong. Someone swaps the embedding model. Still wrong. Someone changes the chunk size, adds a…
Read report
A citation tells you where the system looked. It does not tell you the system was right.
Read report
You built an agent. It worked once. Then you changed the input slightly and it fell apart, and you could not tell which part of the system failed.
Read report
A context window is a bigger desk, not a better memory. The desk still has to be loaded.
Read report
A knowledge base outgrows the prompt, and two camps start shouting. One says the context window is finally big enough, so stop building retrieval…
Read report
The attacker never talks to your model. They leave a sentence in a document it will read later.
Read report
A wrong RAG answer is rarely a model problem. It is a pipeline problem wearing a model costume.
Read report
A demo runs on a clean prompt, a tidy retrieval index, and a cooperative user who types exactly what the script expects. Then the system ships, and inputs…
Read report
A retrieval system is a context-selection system. The model has a limited desk, and every irrelevant chunk takes space away from the evidence it actually…
Read report
A model can name every object in a video and still get the story wrong. That gap is the whole problem.
Read report