Can a small model replace the big LLM that labels evaluation reports?

Yes, with a big it depends. We scored 55 variants of 9 open models for EvalExplorer. A fine-tuned 2-billion-parameter model, about 60 times smaller than the pipeline's 117-billion-parameter LLM (gpt-oss-120b), gives the same labels on 85% of fields on average, and as a 4-bit file fits on a laptop. On the common labels it is reliable. On rare and loosely defined ones, no model we tried does well, and the reason is the training data, not the model.

Training the model was the quick part. This started as a 2-hour internal hackathon at Baobab Tech. The results show where the real work is: a precise codebook, enough verified examples of every label, and a test set that can measure each one. That is data preparation, and it is worth not rushing.

Frugal by design. Training the 2B model is one GPU for 21 minutes, under $1. Labelling then takes 0.4 to 1 second per report on one A100, a single GPU instead of a large hosted model. Small enough to run on your own machine: on a MacBook Pro (M5 Max), the 350M model as an 8-bit GGUF labelled the 134 test reports in 55 seconds (score 79.8). The 2B and 4B files have not been timed on a laptop yet, and energy use was not measured.

~60×fewer parameters than the pipeline's LLM (2B against 117B)
84.7agreement with the pipeline, best model (Qwen3.5 2B), out of 100
2.8 GBQwen3.5-4B as a 4-bit GGUF file, scoring 84.1
$0.89to fine-tune Qwen3.5-2B (21 min on one A100), scoring 84.2
$3.20for the top model: that fine-tune plus GRPO, 77 min in all

The question

When an evaluation report enters EvalExplorer, the ingestion pipeline sends its first pages to a large LLM (gpt-oss-120b, with Gemini 2.5 Flash and Qwen 3 235B as fallbacks). It returns five labels: the evaluation approach (mixed methods, experimental, ...), its type (impact evaluation, systematic review, ...), its timing (baseline, midterm, endline), its themes (global health, governance, ...) and the countries it covers.

How small can a model be and still give the same answers, so that this runs on a laptop or cheaply at scale, without calling a big LLM for every report?

The labels

Each report gets five fields. The codes and definitions below are the ones every model was given; the bars show how often the pipeline used each code across the 1,420 reports.

Approach evaluation_approach

One code, or blank when the report does not say.

  • mixed_methodsCombination of quantitative and qualitative approaches44%
  • experimentalRandomized experiments with control groups25%
  • theory_basedTheory of change or logic model driven15%
  • quasi_experimentalNon-randomized comparison groups10%
  • participatoryStakeholder involvement in evaluation design4%
  • developmentalIterative, adaptive evaluation approach2%

Type evaluation_type

One code, or blank.

  • impact_evaluationExamines the changes caused by an intervention52%
  • process_evaluationExamines activities in an intervention's implementation and the pathways by which the policy was delivered30%
  • systematic_reviewType of synthesis review that employs repeatable methods to find, select and synthesize available evidence on a specific research question11%
  • blankno code given5%
  • rapid_evidence_assessmentSystematic but rapid literature reviews3%

Timing temporality

One code, or blank.

  • endlineAfter completion/at end of program60%
  • blankno code given19%
  • midtermDuring implementation15%
  • baselineBefore intervention/program start6%

Themes themes

One to four codes. Reports carry 2.8 themes on average.

  • social_developmentPoverty, inequality, social protection56%
  • global_healthHealth systems, disease, WASH, nutrition32%
  • governanceGovernment, institutions, rule of law28%
  • gender_equalitiesGender equality, women's rights, disability, LGBT+27%
  • educationEducation programs and outcomes25%
  • economic_developmentDevelopment finance, infrastructure24%
  • food_agricultureFood security, agriculture, farming18%
  • humanitarianEmergency response, disaster relief17%
  • growthEconomic growth, trade, business9%
  • climateClimate change, adaptation, mitigation7%
  • information_digitalICT, digital development7%
  • conflictPeace, security, conflict prevention5%
  • global_partnershipsInternational cooperation, technical assistance5%
  • science_technologyResearch, innovation, technology transfer4%
  • international_financeDevelopment finance, private sector4%
  • infrastructurePhysical infrastructure, construction4%
  • nature_environmentEnvironmental protection, biodiversity3%
  • civil_societyNGOs, community organizations, civic engagement3%

Countries countries

ISO 3166-1 alpha-2 codes of the countries the evaluation covers; empty if none. 145 countries appear in all, 1.6 per report on average; the twelve most frequent are shown.

  • KEKenya14%
  • INIndia12%
  • UGUganda10%
  • BDBangladesh9%
  • ETEthiopia8%
  • TZTanzania6%
  • NGNigeria6%
  • PKPakistan6%
  • GHGhana5%
  • NPNepal4%
  • MWMalawi4%
  • ZASouth Africa4%

What we did

  1. Took 1,420 reports the pipeline had already labelled: 1,148 to train on, 134 kept aside as the test.
  2. Fine-tuned 9 small open models (350M to 26B parameters) to copy the pipeline's answers, with LoRA, and for two of them reinforcement learning (GRPO) on top.
  3. Scored 55 variants in all (zero-shot baselines, fine-tunes, GRPO variants and GGUF exports) on the 134 test reports: how often does each give the same labels as the pipeline?

Results: best run per model

ModelMethodScoreBefore fine-tuningGainvs 3-LLM majoritySeconds / report
Qwen3.5 2B2BSFT + GRPO, countries reward, lr 5e-684.745.8+38.976.70.67report
Qwen3.5 4B4BSFT84.767.1+17.577.61.22report
Gemma 4 26B-A4B26B, 4B activeSFT84.470.1+14.381.01.46report
Gemma 4 E4B4B effectiveSFT83.072.3+10.776.21.65report
Gemma 4 E2B2B effectiveSFT + GRPO, all-fields reward, lr 5e-682.764.9+17.874.91.22report
LFM2.5 1.2B1.2BSFT80.242.0+38.274.00.50report
LFM2.5 350M350MSFT79.220.9+58.371.20.58report
GLiNER2.5 base194Mfine-tune, one passage58.445.4+13.055.60.03report
GLiNER2.5 small74Mfine-tune, chunks57.348.7+8.552.80.04report

Score: agreement with the pipeline's labels on the 134 test reports, 0 to 100 (per report, 1 or 0 for approach, type and timing, F1 for themes and countries, then the average). Differences under about 3 points are within noise for 134 reports. "vs 3-LLM majority" is explained below. Every run, including the ones not shown, is in the experiments repo.

Run it locally

The strongest adapters exported to GGUF for llama.cpp. The 8-bit files match the original models; 4-bit costs a point or two for the 2B models and almost nothing for Qwen3.5-4B.

ModelOriginal8-bit (Q8_0)4-bit (Q4_K_M)4-bit size
qwen3.5-4b-sft84.784.384.12.8 GBfiles
qwen3.5-2b-grpo-countries84.784.882.81.3 GBfiles
gemma-4-26b-a4b-sft84.479.081.516.8 GBfiles
gemma-4-e2b-grpo-lr5e682.782.180.83.4 GBfiles

Gemma 4 26B-A4B scores lower as GGUF than its original run because its adapter behaves differently in plain transformers than in Unsloth, where it was trained and first scored; details in the GGUF repo.

Train your own

What one model costs on Hugging Face Jobs, measured from the jobs that produced the results above (A100 at $2.50 an hour, H200 at $5). Each job also scores the 134 test reports; the training data is 1,148 labelled reports.

ModelStepsGPUTimeCostScore
Qwen3.5 2BLoRA SFTA10021 min$0.8984.2
Qwen3.5 2B (top)LoRA SFT, then GRPO with a countries rewardA10077 min$3.2084.7
Qwen3.5 4BLoRA SFTA10038 min$1.6084.7
Gemma 4 26B-A4BLoRA SFTH20035 min$2.9084.4
GGUF export of one modelmerge, quantize, score 8 variantsA10023-39 min$1-1.60–

For a similar task of your own: about a thousand labelled examples, the same recipe and scripts (code/jobs/sft.py, grpo.py, gguf.py in the experiments repo), and a few dollars per model. The whole study here, 55 variants of 9 models, cost about $45; the label-quality follow-on added about $26 of LLM relabelling.

Where it works and where it doesn't

The score above is an average over five fields, and it hides where the model fails. Only about 1 report in 4 has all five fields right. Below, every code: how many training examples it had, how far the pipeline and three newer LLMs agree on it (a measure of how well defined it is), and how often Qwen3.5 2B (SFT + GRPO, countries reward, lr 5e-6) finds it on the test set.

CodeStatusFoundLabellers agreeTrainTest
experimentalApproachreliable92%9427838
mixed_methodsApproachmixed90%6550860
quasi_experimentalApproachreliable90%7511910
participatoryApproachrare and loosely defined71%43497
theory_basedApproachloosely defined65%3617817
developmentalApproachrare and loosely defined0%38162
impact_evaluationTypereliable94%9159668
systematic_reviewTypereliable92%8613412
process_evaluationTypereliable86%7533637
rapid_evidence_assessmentTyperare and loosely defined38%54348
blankTyperare and loosely defined0%9489
endlineTimingreliable95%8569480
midtermTimingmixed76%6416125
baselineTimingmixed67%69743
blankTimingmixed58%6721926
information_digitalThemesreliable100%82885
nature_environmentThemesreliable100%85323
humanitarianThemesreliable96%8819426
educationThemesreliable91%9327635
climateThemesreliable90%838310
gender_equalitiesThemesreliable88%8431248
social_developmentThemesmixed86%6364079
conflictThemesreliable80%80605
global_healthThemesreliable80%9237440
governanceThemesmixed78%7731040
economic_developmentThemesloosely defined77%3128726
food_agricultureThemesmixed73%8820826
civil_societyThemesrare and loosely defined50%32274
growthThemesmixed50%621006
infrastructureThemesrare50%67424
science_technologyThemesrare and loosely defined50%53512
international_financeThemesrare and loosely defined25%51464
global_partnershipsThemesrare and loosely defined0%335810

"Found": recall of the model on the test reports. "Labellers agree": F1 between the pipeline's labels and the 2-of-3 majority of GLM-5.3-Flash, DeepSeek-V4.1-Flash and Qwen3.8-2.4T-A95B over all 1,420 reports. Recall is measured on the 134 test reports; with fewer than about 10 test examples ("Test") it is a rough figure. "Train": training examples. Reliable: found at least 80% of the time and labellers agree at least 70%. Rare: under 60 training examples. Loosely defined: labellers agree under 60%.

Why: the labels, not the model

LabellersAgreementApproachThemesCountries
DeepSeek-V4.1-Flash – Qwen3.8-2.4T-A95B88.180.286.598.2
GLM-5.3-Flash – DeepSeek-V4.1-Flash86.879.881.897.5
GLM-5.3-Flash – Qwen3.8-2.4T-A95B85.878.084.896.8
Pipeline – DeepSeek-V4.1-Flash76.363.474.289.6
Pipeline – GLM-5.3-Flash76.062.871.990.4
Pipeline – Qwen3.8-2.4T-A95B73.854.274.889.2

Mean field score between two label sets over all 1,420 reports. The pipeline is a 2025 model; the three relabellers are 2026 models given the same pages and code definitions. Details: the follow-on.

Data preparation is the work

What we would do before relying on the rare and loosely defined labels, in order:

  1. Use the full taxonomy, and fix its overlaps. Bring the complete definitions and keyword lists into the labelling prompt and the model's prompt, resolve the codes whose keyword lists overlap, add an example and a counter-example for each neighbouring pair, and write a rule for when a field is blank.
  2. Have people verify a sample. A few hundred reports, weighted towards the codes that are rare or loosely defined, so there is a gold set to measure against.
  3. Collect enough examples of every code. Keep the rare codes; aim for at least 50 verified training examples each, and a test set with 20 to 30 per code so each one can be measured.
  4. Then retrain. At $1 to $3 per model, this is the cheap step.

Next steps

Read more