Minerva Labs Technical resource · 2026

LLMs on your own corpus

From the fundamentals to an assistant that works over your team's documents, code and data.

Eight modules · one hour No prior ML required Reference appendix for implementation
Framing

What you leave with

1

Technical judgement

What an LLM is and is not — so you neither buy hype nor discard useful tools out of prejudice.

2

A working system

An assistant that answers from your own material with a citation, writes and runs code, and checks what it produces.

3

A roadmap

When retrieval is enough, when fine-tuning earns its cost, and when neither is justified.

The deliverable is not a model. It is a workflow your team can maintain, evaluate and change without depending on anyone outside it.

Framing

Map

01Fundamentalstokens → attention → training
05Local inferencequantization, KV cache, cost
02The complexity laddereight levels, one decision
06Fine-tuningwhat it fixes, what it can't
03Retrievalingestion → citations → eval
07Verifierschecking output automatically
04The harnesswhere the value actually is
08Architecture and roadmapphases, hardware, privacy
Module 01

Fundamentals

What an LLM actually is. Understanding the object is what lets you predict where it fails.

Tokens · embeddings · attention · training · failure modes
01 · Fundamentals

The definition fits on one line

p( x_t | x_1, x_2, …, x_t-1 ; θ )

A probability distribution over the next token, conditioned on everything before it, parameterised by θ — billions of real numbers.

Everything else follows

Translating, summarising, deriving, coding — all restated as "predict the most likely continuation".

There is no database inside

θ stores statistical regularities, not indexed facts. That is why it can misremember with total fluency.

Context is the only state

Nothing persists between calls. What is not in the context window does not exist for the model.

01 · Fundamentals

Step one: text becomes integers

The model never sees characters. A tokenizer splits text into frequent sub-word pieces, each with an index in a vocabulary of 50,000–200,000 entries.

"quantum decoherence" → ["quant", "um", " dec", "oher", "ence"] → [1902, 372, 4102, 21985, 4419]

Notation fragments badly

Formulae, identifiers, part numbers and file paths split into meaningless pieces. The model works harder to reassemble them.

Everything is priced per token

Cost, context window and speed are all measured in tokens. A twelve-page document is roughly 8,000–12,000.

Not all languages cost the same

The same content in Spanish spends 15–30% more tokens than in English on most tokenizers. Budget accordingly.

01 · Fundamentals

Step two: tokens become vectors

Each index maps to a vector in R^d, d ≈ 1,000–8,000. Training places those vectors so that the geometry of the space encodes relations of meaning.

Proximity is similar usage, not synonymy
"buy" and "sell" sit close together: they appear in the same contexts, though they mean opposite things.
The standard measure is cosine
cos(u,v) = ⟨u,v⟩ / (‖u‖‖v‖)
A whole text is also a vector
An embedding model compresses a full paragraph into one vector. This is the central piece of retrieval.

The limits

Individual directions have no interpretation — there is no axis that means "temperature". Only relative distances matter, and only approximately.

The space is learned from text, so it inherits its biases: what appears together stays together, whether or not that is correct.

01 · Fundamentals

Step three: attention

Each token builds three projections — query, key and value — and updates its representation as a weighted mixture of every other token's values.

Attention(Q,K,V) = softmax( Q Kᵀ / √d ) · V

QKᵀ measures affinity

How much each token cares about each other token — an inner product between two linear projections.

Softmax normalises

Those overlaps become positive weights summing to one, giving a convex combination of the values.

Cost is quadratic in length

Doubling the context quadruples the compute. This is why context windows are finite and expensive.

Multi-head: many questions at once

Dozens run in parallel with different projections; some track syntax, others long-range reference.

01 · Fundamentals

Step four: stack the block, and pick a size

Attention
Mixes information between positions
MLP
Transforms each position separately
Residual
x + f(x) preserves the original signal
LayerNorm
Stabilises scales between layers

One block does nothing interesting. Capability emerges from stacking 30–120 identical blocks and training them together.

4–9B

Small models. Run on any laptop with a GPU. Useful for auxiliary tasks, weak as an agent.

27–35B

The practical band for a team: one GPU, high quality, open weights. The default choice.

>1T

Frontier models. Some with open weights, but they demand multi-GPU infrastructure.

01 · Fundamentals

How it is trained I: pretraining

The task

Hide the next token and ask the model to predict it. The loss is cross-entropy, the fit is stochastic gradient descent. No human labels: the text supervises itself.

The scale

Ten to thirty trillion tokens — effectively the usable web, books, code and scientific articles. Months of compute on tens of thousands of GPUs. Hundreds of millions of dollars.

The public literature in your field is very probably already inside the model. What is not inside is your unpublished material, your internal notes, your private code and your own data. That gap is exactly what the rest of this deck fills.

Nobody pretrains from scratch outside a compute centre. You always start from open weights already trained — that is the normal starting point, not a limitation.

01 · Fundamentals

How it is trained II: post-training

A pretrained model only continues text. Turning it into a useful assistant takes three further stages, all far cheaper than the first.

SFT
Supervised fine-tuning

Thousands of instruction-and-ideal-answer examples written by humans. Teaches the conversation format and how to follow directions.

RLHF / DPO
Preference learning

Annotators compare pairs of answers. The model learns what is preferred — usefulness, honesty, tone. It does not learn new facts.

RLVR
Verifiable rewards

Training on problems whose answer can be checked automatically — mathematics, code. This is the origin of today's reasoning models, and module 07 returns to it.

01 · Fundamentals

The four failures you will meet

Hallucination

Invents references, identifiers, equations and results with total fluency. No internal signal separates recalled from fabricated.

Mitigation. Force citation of a retrieved source; treat an uncited answer as a failure.

Frozen knowledge

Training has a cut-off date. It does not know last week's preprint or your internal results.

Mitigation. Retrieval over a corpus you keep current; web search for what is public.

Prompt sensitivity

Rephrasing the same question changes the answer. It is not deterministic over meaning.

Mitigation. Fixed, versioned prompts; a stable evaluation set.

Sycophancy

Tends to agree with the user's premise, even when the premise is false.

Mitigation. Ask explicitly for counterarguments; keep your desired conclusion out of the question.

Module 02

The complexity ladder

Eight levels of solution ordered by effort. The most important decision in the project is which rung to stop on.

Effort · value · the question that decides everything
02 · Ladder

Eight levels, simplest to most costly

0A well-written promptA capable model and precise instructions.Hours
1Long contextPaste the relevant documents into the window.Hours
2An open harness, as it shipsInstall it and point it at the corpus.Days
3The harness adapted to your domainSource-format parsing, structural chunking, symbol-aware lexical search.Weeks
4An agent that executes and checksRuns code in isolation and verifies its own output.Weeks
5Fine-tuning with LoRAFix format and house conventions on your own material.1–2 months
6Continued pretrainingKeep training on the whole corpus.Months
7RL with verifiable rewardsTrain against automatic verifiers for your domain.Research line

The levels are not exclusive: a good system combines 0–4. Levels 5–7 sit on top, never instead, and only once the harness is exhausted.

02 · Ladder

Is the problem missing knowledge, or missing behaviour?

Missing knowledge

The model does not know what your 2024 report says, which sign convention your team uses, or what came out of the last experiment.

→ Retrieval. Fine-tuning is the wrong tool here: memorising facts by adjusting weights is expensive, unreliable and impossible to update.

Missing behaviour

The model knows the subject but answers in the wrong format, in the wrong register, or without following the structure of your reports.

→ Fine-tuning. This is where it pays: hundreds or thousands of examples are enough to fix format, style and internal convention.

A third possibility, the most frequent of all: the model knows, but the harness never found the data, never ordered it properly, or never let it run anything to check.

02 · Ladder

Effort against value delivered

Value delivered Effort required
N0N1N2N3N4N5N6N7

How to read it

Level 2 is nearly free: adopting a harness someone else built already delivers a good share of the value, in days.

Between 3 and 4 you reach about 90% of the value for less than half the total effort. That work is pipeline work, not model work.

From 5 onwards the cost multiplies and the improvement is marginal.

Illustrative scales based on comparable projects, not a measurement on your corpus. Measuring it is part of the work.

Module 03

Retrieval

Where your private corpus becomes something queryable, verifiable and updatable.

Ingestion · hybrid search · citations · evaluation
03 · Retrieval

Retrieval in one sentence

Before answering, search the corpus for the relevant fragments and paste them into the context of the question.

A · Indexing — once, offline
Documents, notes, code
Parsing
Chunking
Embeddings
Vector + lexical index
B · Query — every time someone asks
Question
Retrieval
Reranking
Prompt + context
Answer with citations

The decisive advantage over fine-tuning: you update it by adding a file, the answer points at a verifiable source, and nothing forces a retrain when the corpus grows.

03 · Retrieval

Ingestion: the most underestimated problem

A technical PDF is not text. It is a print format with equations as glyphs, figures, tables, two-column flow and footnotes that break reading order.

Use the source format whenever it existsLaTeX, Markdown, DOCX, the original spreadsheet. Structure and notation survive intact — worth more than any PDF extractor.
The corpus is not only documentsCode, notebooks, protocols, instrument configs, result logs, tickets. All indexable text, and the most consulted day to day.
Only if you must: extractorsMarker, Docling or MinerU convert PDF to Markdown with equations in LaTeX. Good, not perfect. Review a sample by hand.
Figures and tables: describe themPass them through a vision model and store the description as indexable text, next to the original caption.
Keep provenance from minute oneEvery fragment with file, identifier, section, page and equation number. Without this there are no citations, and without citations there is no trust.

Rule: final system quality is bounded by ingestion quality. No model compensates for bad parsing.

03 · Retrieval

Hybrid retrieval: the critical piece

Dense search

Compares embeddings by cosine.

Strong. Finds paraphrase and related concepts even with no shared words.

Weak. Fails on symbols, acronyms and proper names.

Lexical search (BM25)

Counts exact term matches.

Strong. Finds a part number, a sample code, an author name, a specific unit.

Weak. Understands neither synonyms nor rephrasing.

Combine them with reciprocal rank fusion: each document contributes 1/(k + rank) from each list, and you re-sort by the sum. No weights to calibrate, no scales to normalise.

score(d) = Σ_lists 1 / (60 + rank_list(d))   # k = 60 works with no tuning
03 · Retrieval

Generating with mandatory citations

Answer only from the fragments provided. Rules: 1. Every claim carries its citation: [doc_id, sec.] 2. If the fragments are not enough, say so and stop. 3. Do not fill gaps with general knowledge. 4. Distinguish measured from conjectured. 5. Reproduce formulae literally. <fragments> {retrieved_context} </fragments> Question: {question}
Verifiable citations
The reader must be able to open the document and check the sentence. Without that, a technical team will not adopt it.
Explicit permission not to know
"There is not enough information" has to be an acceptable and frequent answer.
Separate retrieved from generated
In the interface, show the fragments used next to the answer.
Low temperature
0 to 0.2. Creativity is not what you are after here.

An answer with no retrieved citation is a system failure, not a weaker answer.

03 · Retrieval

Evaluation: without it you are blind

Before writing a line of code: 50–100 cases from your own work with known answers. Questions with their source, and tasks whose result can be checked by running them.

Recall@k

In what fraction of questions does the correct fragment appear among the k retrieved? Measures the retriever, isolated from the generator.

Faithfulness

Is every claim in the answer supported by a retrieved fragment? Can be automated with a judge model.

Citation precision

Does the citation point at the right place? This is what destroys trust fastest when it fails.

Tasks solved

Of the executable tasks in the set, in how many does the produced result match the known one? The metric hardest to fake.

Module 04

The harness is the project

Almost all the real complexity lives here — not in the model, not in training, but in the scaffolding around it.

Adopt · adapt · execute · only then train
04 · Harness

The argument that orders the project

A model is worth what its harness is worth. But a harness cannot be taught new knowledge — only training does that. The correct order of work follows.

1
Adopt an existing harness
Open projects exist that are built exactly for querying a document corpus with citations. Installing one and pointing it at your files is days, not weeks.
2
Adapt it, in Python, to your domain
This is the bulk of the value. Generic harnesses are good, but they return far more once parsing, chunking, search and prompts are fitted to your own material.
3
Give it execution and checking
Let it run code in an isolated environment and verify its own output before delivering. Requires no training and raises reliability measurably.
4
Only then, fine-tuning
When the harness is exhausted and evaluation points at a limit the scaffolding cannot resolve.

Published evidence: changing only the harness, with the same model, moves results on agentic benchmarks more than changing the model.

04 · Harness

Two layers, and what exists in each

Retrieving and executing are different problems, and mature open projects exist for each. Picking only one leaves half the work uncovered.

Retrieval

Answers questions about the corpus with verifiable citations.

PaperQA2 — a RAG agent for document literature: iterative search, reranking, in-text citations.
RAGFlow — deep document parsing: multi-column, tables and scans.
Execution

Reads files, writes code, runs it, reads the error and tries again.

OpenHands — autonomous agent with a containerised shell by default; runs unattended, no UI required.
OpenCode, Goose, Aider — terminal agents, provider-agnostic, with local-model support.
Pydantic AI — composable file, shell, memory and sub-agent capabilities if you would rather build the loop yourself.

The correct composition: the execution loop is the main one, and retrieval is exposed inside it as one more tool.

04 · Harness

The loop and its tools

The harness is a loop: the model picks a tool, sees the result and decides the next step. What defines its quality is which tools it has — and how few.

1
Retrieve
Search the corpus and return fragments with provenance.
2
Read and write
Access to the file tree: code, notebooks, configs, data.
3
Execute
Python and shell inside an isolated container with a pinned environment.
4
Check
Symbolic algebra, units, limit cases, comparison against known results.
5
Log
Store the full trace: commands, seeds, versions and outputs.

Fewer tools, better results

Cutting an agent's tool set can raise its success rate noticeably while lowering token spend and latency. An agent with twenty tools chooses badly.

Practical corollary: start with these five, measure, and add a sixth only if the evaluation asks for it.

04 · Harness

Where adapting it pays

A generic harness assumes generic documents. Each of these is a small code change with a large, measurable effect on a technical corpus.

Parsing
Read the source format instead of extracting from the PDF wherever it exists.
Chunking
Cut by section and by equation or code block, not at fixed length.
Lexical search
A tokenizer that preserves symbols, acronyms and internal identifiers.
Metadata
Extract author, date, subject, method and publication status at index time.
Prompts
Citation rules, permission to abstain, code and formula formatting.
Tools
Isolated execution, symbolic algebra and dimensional checking inside the loop.

None of this requires knowing how to train models. All of it requires knowing the corpus — which is exactly what your team has and an outside vendor does not.

Module 05

Local inference

Why running it yourself is the default for a team with private material, and what it takes to fit the hardware.

Privacy · reproducibility · quantization · cost
05 · Local

Four reasons to run locally

01

Privacy without conditions

Unpublished material, proposals and internal data never leave your network. No retention terms to review, no institutional agreement to negotiate.

02

Reproducibility

A local checkpoint has a hash: pin it, record it with your results, run it again in two years. An API model changes under your feet without notice.

03

Marginal cost near zero

Once the GPU is bought, each call is electricity. That is what makes long agentic loops, with many attempts and retries, viable at all.

04

No rate or context limits

Reindex the whole corpus, sweep parameters, launch a thousand evaluations overnight without watching a bill.

05 · Local

Weight quantization: how it fits on the card

Weights are stored with fewer bits than they were trained in. Quality drops a little; memory drops a lot. This is what turns "I need a server" into "it fits on one card".

BF16
2 bytes per parameter
≈54 GB

Quality reference for a 27B model. Does not fit a consumer GPU.

FP8
1 byte per parameter
≈27 GB

Almost no loss. Real speed-up on Hopper and later.

4-bit
0.5 bytes per parameter
≈14 GB

GGUF, AWQ or NVFP4. Small loss with calibrated quantization.

Not all layers quantize equally. Dynamic schemes keep sensitive layers at higher precision and calibrate the rest with real data. The difference against uniform quantization is noticeable.

Formats are hardware-bound. NVFP4 needs Blackwell; on earlier cards GGUF and AWQ remain the right path. Check the format before buying hardware, not after.

05 · Local

KV cache and mixture of experts

KV cache

To avoid recomputing attention for every generated token, the keys and values of the whole context are stored. With long windows that cache ends up larger than the weights themselves.

Quantising it to FP8 halves it: twice the context, or twice the users, on the same hardware.

Quality holds well at FP8. Below that, validate against your evaluation set before using it.

Mixture of experts

The model holds many specialised blocks but activates only a few per token. A 35B-A3B stores 35 billion parameters and computes with 3 billion.

Consequence: total parameters set the memory, active parameters set the speed. It suits agentic pipelines, which make many short chained calls.

It suits a memory-tight GPU badly — a smaller dense model will do better there.

Sizing rule: memory ≈ quantized weights + KV cache at the context you will really use + a margin for activations.

05 · Local

Cumulative cost: local against API

Twelve-month cumulative cost, ten people, daily agentic use
Frontier API≈ 18,000 USD
Economy API≈ 7,000 USD
Local GPU + electricity≈ 6,600 USD

How to read it

Against a frontier API the GPU pays for itself around the fourth or fifth month of intensive team use.

Against an economy API the crossing takes much longer. There the argument is not price: it is privacy, reproducibility and being able to spend tokens without thinking.

And token spend matters: a retrieval agent can consume hundreds of thousands of tokens on a single question.

Scenario: ten people, daily use, agentic pipeline, a 6,000 USD GPU amortised at once. Adjust the numbers to your case before quoting them.

Module 06

Fine-tuning

The last resort, not the first. It earns its place only once the harness is adapted, measured and exhausted.

Behaviour, not knowledge · LoRA
06 · Fine-tuning

What it fixes, and what it does not

It does work for

Stable output format — the agent always returns the structure your pipeline expects.
House conventions: your notation, identifiers, units, internal APIs.
Narrow repetitive tasks: classify, extract parameters, normalise tables.
Cutting cost — a tuned small model can match a large one on one specific task.

It does not work for

Memorising specific facts from your documents in a reliably recoverable way.
Keeping knowledge current — every new document would demand a retrain.
Giving verifiable citations: weights do not store provenance.
Fixing hallucination. It usually worsens it: the model sounds more confident about your domain.

Rule: fine-tuning teaches behaviour, the harness teaches searching, knowledge is retrieved. Each in its own layer.

06 · Fine-tuning

LoRA: the trick that makes it viable

Adjusting all the billions of parameters in W is out of reach outside a compute centre. LoRA freezes W and learns a low-rank correction instead.

W' = W + B·A   with r ≈ 8–64 ≪ d, k

Only B and A are trained

Between 0.1% and 2% of the original parameters. The rest stays frozen.

It fits on one GPU

The same card that serves inference can train the adapter, usually overnight.

Adapters are swappable

One base model, several small adapters — one per task, versioned like code.

Module 07

Verifiers

Automatic checking is valuable on its own, and it is the door to training with verifiable rewards.

What you can check · four phases
07 · Verifiers

Verifiers that already exist in your domain

Symbolic algebraChecks whether two expressions are equivalent, whether a derivation is valid, whether a limit matches.High
Unit and dimensional consistencyEvery proposed equation must be dimensionally correct. Verifiable in milliseconds, always available.High
Tests and type checks on codeThe cheapest verifier of all, and the one most teams already have sitting in the repository.High
Known limit casesCheck that a result reproduces the classical limit, the non-interacting case, the already-published regime.Medium
Numerical simulationA solver evaluates whether the analytic prediction matches simulation within tolerance.Medium
Your own recorded resultsThe figures in your tables and plots are checkable. An honest test set, if a small one.Medium

Verifier reliability →

07 · Verifiers

A research line in four phases

Structured so each phase delivers something usable on its own. You can stop, pause or rethink at any point without losing the earlier work.

V1
Standalone verifiers
Dimensional checking, symbolic equivalence, limit cases. Used by hand on the assistant's answers. Already catches real errors.
V2
Verifiers inside the agent loop
The agent checks its own answer before delivering and retries on failure. Nothing is trained. Reliability rises measurably.
V3
A bench of verifiable problems
A set of domain problems with automatic verification. A publishable artefact in its own right, and a citable benchmark.
V4
Training with RLVR
Only with V3 standing: train against the verifiers and measure whether the model improves where it matters.

V1 and V2 are part of the harness and fit a normal project calendar. V3 and V4 are research, with their own.

Module 08

Architecture and roadmap

A concrete architecture, a realistic calendar, and a system that runs entirely inside your own network.

Four layers · eight weeks · one GPU
08 · Architecture

Target architecture

Corpus
Source files and parsed PDFs
Drafts and internal notes
Code, notebooks, configs
Logs of previous results
Metadata per fragment
Retrieval
Dense index + BM25 lexical
Reranking and contextual summary
Iterative search
Metadata filters
Exposed as a single tool
Execution
Python and shell in a container
Environment and seeds pinned
Symbolic and unit checking
Retry when a check fails
Full trace recorded
Local inference
Open weights served with vLLM
Weights at 4-bit or FP8
KV cache in FP8
Local embeddings and reranker
Nothing leaves the network

Levels 2 to 4 of the ladder. No fine-tuning in the initial version: it is added later, and only if the evaluation asks for it.

08 · Architecture

Roadmap by phase

Days 1–3Evaluation set and corpusA golden set of questions and an inventory of the corpus. The one step you cannot skip.
Week 1Open harness + local modelThe harness over your document folder with an open model behind it. It already answers with citations.
Weeks 2–4Adaptation in PythonSource-format parsing, structural chunking, symbol-aware lexical search, prompts. The bulk of the value.
Weeks 5–8Execution and checkingAn isolated execution loop with retrieval as a tool. Symbolic checks, units, trace logging.
OptionalLoRA on your own materialOnly if the evaluation points at a limit the harness cannot resolve.
ResearchVerifiable bench and RLVRIts own calendar and people, with a decision point at the end of each phase.

Eight weeks to a system in real use. Most of that time is pipeline adaptation, not training or infrastructure.

08 · Architecture

Realistic effort and cost

1
GPU of 24–48 GB

One-off cost. Serves inference, embeddings, reranking — and later fine-tuning too.

~0
Per query

After the hardware, each question is electricity. That is what allows agentic pipelines without a brake.

1
Person, half-time

Someone fluent in Python. No prior machine-learning experience required.

8
Weeks to real use

With something answering with citations already in week one.

Useful comparison: one serious attempt at fine-tuning costs more person-time than phases 0 to 3 combined, and produces less observable value for the team.

Action

The eight mistakes you are going to make

1Adapting code in week one, before seeing the pieces work untouched.
5Ignoring parse quality until someone notices the equations are garbage.
2Reimplementing from scratch a harness that already exists and is built for this.
6Changing three things at once and not knowing which moved the number.
3Sizing the GPU from the weights alone and forgetting the KV cache at long context.
7Giving the agent twenty tools and being surprised it chooses badly.
4Never building the case set, then deciding by feel whether the system improved.
8Letting it execute with no isolation, no pinned environment and no trace.

They are listed in the order they appear, week by week. The most expensive is the fourth: without metrics, the others are invisible.

Three ideas to take away
1

A model is worth what its harness is worth.

Retrieving, executing and checking are the scaffolding, and nearly all the value sits there. Adopt one that exists, adapt it, and only then train.

2

Local by default.

Privacy without conditions, citable reproducibility, and a marginal cost near zero — which is what makes an agentic pipeline viable in the first place.

3

Without evaluation there is no project.

A hundred cases with known answers, executable wherever possible, are worth more than any architecture decision.

Next step: assemble the evaluation set and stand up the harness. One week, and the rest of the project stops being a bet.

Minerva Labs

Questions, doubts and disagreements

Written as a working resource. Take it, adapt it to your corpus, and tell us where it was wrong.

contact@minervalabs.mx Minerva Labs · edge-first AI research