Minerva Labs
Technical resource · 2026
From the fundamentals to an assistant that works over your team's documents, code and data.
What an LLM is and is not — so you neither buy hype nor discard useful tools out of prejudice.
An assistant that answers from your own material with a citation, writes and runs code, and checks what it produces.
When retrieval is enough, when fine-tuning earns its cost, and when neither is justified.
The deliverable is not a model. It is a workflow your team can maintain, evaluate and change without depending on anyone outside it.
What an LLM actually is. Understanding the object is what lets you predict where it fails.
A probability distribution over the next token, conditioned on everything before it, parameterised by θ — billions of real numbers.
Translating, summarising, deriving, coding — all restated as "predict the most likely continuation".
θ stores statistical regularities, not indexed facts. That is why it can misremember with total fluency.
Nothing persists between calls. What is not in the context window does not exist for the model.
The model never sees characters. A tokenizer splits text into frequent sub-word pieces, each with an index in a vocabulary of 50,000–200,000 entries.
Formulae, identifiers, part numbers and file paths split into meaningless pieces. The model works harder to reassemble them.
Cost, context window and speed are all measured in tokens. A twelve-page document is roughly 8,000–12,000.
The same content in Spanish spends 15–30% more tokens than in English on most tokenizers. Budget accordingly.
Each index maps to a vector in R^d, d ≈ 1,000–8,000. Training places those vectors so that the geometry of the space encodes relations of meaning.
Individual directions have no interpretation — there is no axis that means "temperature". Only relative distances matter, and only approximately.
The space is learned from text, so it inherits its biases: what appears together stays together, whether or not that is correct.
Each token builds three projections — query, key and value — and updates its representation as a weighted mixture of every other token's values.
How much each token cares about each other token — an inner product between two linear projections.
Those overlaps become positive weights summing to one, giving a convex combination of the values.
Doubling the context quadruples the compute. This is why context windows are finite and expensive.
Dozens run in parallel with different projections; some track syntax, others long-range reference.
One block does nothing interesting. Capability emerges from stacking 30–120 identical blocks and training them together.
Small models. Run on any laptop with a GPU. Useful for auxiliary tasks, weak as an agent.
The practical band for a team: one GPU, high quality, open weights. The default choice.
Frontier models. Some with open weights, but they demand multi-GPU infrastructure.
Hide the next token and ask the model to predict it. The loss is cross-entropy, the fit is stochastic gradient descent. No human labels: the text supervises itself.
Ten to thirty trillion tokens — effectively the usable web, books, code and scientific articles. Months of compute on tens of thousands of GPUs. Hundreds of millions of dollars.
The public literature in your field is very probably already inside the model. What is not inside is your unpublished material, your internal notes, your private code and your own data. That gap is exactly what the rest of this deck fills.
Nobody pretrains from scratch outside a compute centre. You always start from open weights already trained — that is the normal starting point, not a limitation.
A pretrained model only continues text. Turning it into a useful assistant takes three further stages, all far cheaper than the first.
Thousands of instruction-and-ideal-answer examples written by humans. Teaches the conversation format and how to follow directions.
Annotators compare pairs of answers. The model learns what is preferred — usefulness, honesty, tone. It does not learn new facts.
Training on problems whose answer can be checked automatically — mathematics, code. This is the origin of today's reasoning models, and module 07 returns to it.
Invents references, identifiers, equations and results with total fluency. No internal signal separates recalled from fabricated.
Mitigation. Force citation of a retrieved source; treat an uncited answer as a failure.
Training has a cut-off date. It does not know last week's preprint or your internal results.
Mitigation. Retrieval over a corpus you keep current; web search for what is public.
Rephrasing the same question changes the answer. It is not deterministic over meaning.
Mitigation. Fixed, versioned prompts; a stable evaluation set.
Tends to agree with the user's premise, even when the premise is false.
Mitigation. Ask explicitly for counterarguments; keep your desired conclusion out of the question.
Eight levels of solution ordered by effort. The most important decision in the project is which rung to stop on.
The levels are not exclusive: a good system combines 0–4. Levels 5–7 sit on top, never instead, and only once the harness is exhausted.
The model does not know what your 2024 report says, which sign convention your team uses, or what came out of the last experiment.
→ Retrieval. Fine-tuning is the wrong tool here: memorising facts by adjusting weights is expensive, unreliable and impossible to update.
The model knows the subject but answers in the wrong format, in the wrong register, or without following the structure of your reports.
→ Fine-tuning. This is where it pays: hundreds or thousands of examples are enough to fix format, style and internal convention.
A third possibility, the most frequent of all: the model knows, but the harness never found the data, never ordered it properly, or never let it run anything to check.
Level 2 is nearly free: adopting a harness someone else built already delivers a good share of the value, in days.
Between 3 and 4 you reach about 90% of the value for less than half the total effort. That work is pipeline work, not model work.
From 5 onwards the cost multiplies and the improvement is marginal.
Illustrative scales based on comparable projects, not a measurement on your corpus. Measuring it is part of the work.
Where your private corpus becomes something queryable, verifiable and updatable.
Before answering, search the corpus for the relevant fragments and paste them into the context of the question.
The decisive advantage over fine-tuning: you update it by adding a file, the answer points at a verifiable source, and nothing forces a retrain when the corpus grows.
A technical PDF is not text. It is a print format with equations as glyphs, figures, tables, two-column flow and footnotes that break reading order.
Rule: final system quality is bounded by ingestion quality. No model compensates for bad parsing.
Compares embeddings by cosine.
Strong. Finds paraphrase and related concepts even with no shared words.
Weak. Fails on symbols, acronyms and proper names.
Counts exact term matches.
Strong. Finds a part number, a sample code, an author name, a specific unit.
Weak. Understands neither synonyms nor rephrasing.
Combine them with reciprocal rank fusion: each document contributes 1/(k + rank) from each list, and you re-sort by the sum. No weights to calibrate, no scales to normalise.
An answer with no retrieved citation is a system failure, not a weaker answer.
Before writing a line of code: 50–100 cases from your own work with known answers. Questions with their source, and tasks whose result can be checked by running them.
In what fraction of questions does the correct fragment appear among the k retrieved? Measures the retriever, isolated from the generator.
Is every claim in the answer supported by a retrieved fragment? Can be automated with a judge model.
Does the citation point at the right place? This is what destroys trust fastest when it fails.
Of the executable tasks in the set, in how many does the produced result match the known one? The metric hardest to fake.
Almost all the real complexity lives here — not in the model, not in training, but in the scaffolding around it.
A model is worth what its harness is worth. But a harness cannot be taught new knowledge — only training does that. The correct order of work follows.
Published evidence: changing only the harness, with the same model, moves results on agentic benchmarks more than changing the model.
Retrieving and executing are different problems, and mature open projects exist for each. Picking only one leaves half the work uncovered.
Answers questions about the corpus with verifiable citations.
Reads files, writes code, runs it, reads the error and tries again.
The correct composition: the execution loop is the main one, and retrieval is exposed inside it as one more tool.
The harness is a loop: the model picks a tool, sees the result and decides the next step. What defines its quality is which tools it has — and how few.
Cutting an agent's tool set can raise its success rate noticeably while lowering token spend and latency. An agent with twenty tools chooses badly.
Practical corollary: start with these five, measure, and add a sixth only if the evaluation asks for it.
A generic harness assumes generic documents. Each of these is a small code change with a large, measurable effect on a technical corpus.
None of this requires knowing how to train models. All of it requires knowing the corpus — which is exactly what your team has and an outside vendor does not.
Why running it yourself is the default for a team with private material, and what it takes to fit the hardware.
Unpublished material, proposals and internal data never leave your network. No retention terms to review, no institutional agreement to negotiate.
A local checkpoint has a hash: pin it, record it with your results, run it again in two years. An API model changes under your feet without notice.
Once the GPU is bought, each call is electricity. That is what makes long agentic loops, with many attempts and retries, viable at all.
Reindex the whole corpus, sweep parameters, launch a thousand evaluations overnight without watching a bill.
Weights are stored with fewer bits than they were trained in. Quality drops a little; memory drops a lot. This is what turns "I need a server" into "it fits on one card".
Quality reference for a 27B model. Does not fit a consumer GPU.
Almost no loss. Real speed-up on Hopper and later.
GGUF, AWQ or NVFP4. Small loss with calibrated quantization.
Not all layers quantize equally. Dynamic schemes keep sensitive layers at higher precision and calibrate the rest with real data. The difference against uniform quantization is noticeable.
Formats are hardware-bound. NVFP4 needs Blackwell; on earlier cards GGUF and AWQ remain the right path. Check the format before buying hardware, not after.
To avoid recomputing attention for every generated token, the keys and values of the whole context are stored. With long windows that cache ends up larger than the weights themselves.
Quantising it to FP8 halves it: twice the context, or twice the users, on the same hardware.
Quality holds well at FP8. Below that, validate against your evaluation set before using it.
The model holds many specialised blocks but activates only a few per token. A 35B-A3B stores 35 billion parameters and computes with 3 billion.
Consequence: total parameters set the memory, active parameters set the speed. It suits agentic pipelines, which make many short chained calls.
It suits a memory-tight GPU badly — a smaller dense model will do better there.
Sizing rule: memory ≈ quantized weights + KV cache at the context you will really use + a margin for activations.
Against a frontier API the GPU pays for itself around the fourth or fifth month of intensive team use.
Against an economy API the crossing takes much longer. There the argument is not price: it is privacy, reproducibility and being able to spend tokens without thinking.
And token spend matters: a retrieval agent can consume hundreds of thousands of tokens on a single question.
Scenario: ten people, daily use, agentic pipeline, a 6,000 USD GPU amortised at once. Adjust the numbers to your case before quoting them.
The last resort, not the first. It earns its place only once the harness is adapted, measured and exhausted.
Rule: fine-tuning teaches behaviour, the harness teaches searching, knowledge is retrieved. Each in its own layer.
Adjusting all the billions of parameters in W is out of reach outside a compute centre. LoRA freezes W and learns a low-rank correction instead.
Between 0.1% and 2% of the original parameters. The rest stays frozen.
The same card that serves inference can train the adapter, usually overnight.
One base model, several small adapters — one per task, versioned like code.
Automatic checking is valuable on its own, and it is the door to training with verifiable rewards.
Verifier reliability →
Structured so each phase delivers something usable on its own. You can stop, pause or rethink at any point without losing the earlier work.
V1 and V2 are part of the harness and fit a normal project calendar. V3 and V4 are research, with their own.
A concrete architecture, a realistic calendar, and a system that runs entirely inside your own network.
Levels 2 to 4 of the ladder. No fine-tuning in the initial version: it is added later, and only if the evaluation asks for it.
Eight weeks to a system in real use. Most of that time is pipeline adaptation, not training or infrastructure.
One-off cost. Serves inference, embeddings, reranking — and later fine-tuning too.
After the hardware, each question is electricity. That is what allows agentic pipelines without a brake.
Someone fluent in Python. No prior machine-learning experience required.
With something answering with citations already in week one.
Useful comparison: one serious attempt at fine-tuning costs more person-time than phases 0 to 3 combined, and produces less observable value for the team.
They are listed in the order they appear, week by week. The most expensive is the fourth: without metrics, the others are invisible.
Retrieving, executing and checking are the scaffolding, and nearly all the value sits there. Adopt one that exists, adapt it, and only then train.
Privacy without conditions, citable reproducibility, and a marginal cost near zero — which is what makes an agentic pipeline viable in the first place.
A hundred cases with known answers, executable wherever possible, are worth more than any architecture decision.
Next step: assemble the evaluation set and stand up the harness. One week, and the rest of the project stops being a bet.
Minerva Labs
Written as a working resource. Take it, adapt it to your corpus, and tell us where it was wrong.