# The role of the language model

The Devseis Endpoint Auditor uses a small language model (Qwen2.5-0.5B-Instruct, fine-tuned by Devseis, running on the
device with WebLLM). This document says what the model does, what it deliberately does not do, how its output is
checked, and how its role can grow.

**In one sentence:** code collects the evidence and decides every verdict; the model explains, prioritises and writes
the report, and is checked against the evidence before anything it writes is used.

## Who does what

| Job | Done by | Where | Why |
|---|---|---|---|
| Read the device's settings (password policy, encryption, firewall, …) | Collector scripts (read-only) | `collectors/` | A model cannot read a computer; read-only scripts can, and change nothing |
| Phones: detect OS, version, screen lock; ask the rest | Browser detection + guided questions | `mobile/detect.js`, `catalog/mobile_questions.json` | A web page cannot read phone settings; answers are labelled self-reported |
| Decide Compliant / Non-Compliant / Pending / NotApplicable / Error | Rules in code, from `baseline.conf` | `collectors/`, `mobile/mobile-core.js` | A compliance verdict must be exact, repeatable and explainable |
| Match installed software against vulnerabilities | Vulnerability matcher (exact dpkg / rpm / version-range comparison) | `app/renderer/vuln-core.js`, `vulndb/` | A model must never guess whether a CVE applies |
| Score, counts, top gaps | Code | `app/renderer/auditor-core.js` (`summaryTemplate`) | Numbers in the report always come from the evidence |
| **Explain each finding** in plain language | **Model** | `findingPrompt` → model | Turns "FileVault recovery key missing" into what it means and why it matters |
| **Assess the risk** of each gap | **Model** (guided by the fix library) | `library/check_guidance.json` | High / Medium / Low with a reason |
| **Write fix and verification steps** for the device's OS | **Model** (quoting the fix library) | `library/check_guidance.json` | Mac steps for a Mac, Android steps for Android |
| **Map findings to ISO 27001, GDPR and the EU AI Act** | **Model** (from retrieved references) | `library/controls.json`, `catalog/checks.json` | Which control or article applies, in readable words |
| **Write the executive summary and next steps** | **Model** | `summaryPrompt` → model | A manager-level view, in priority order |
| **Classify AI tools under the EU AI Act** | **Model** | `aiPrompt` → model, `library/ai_tools.json` | Risk class, GDPR concerns and an action per tool |
| Say what the evidence cannot prove | **Model** | finding text | e.g. "Self-reported by the user, not verified on the device" |

## How a report is written

```
collector / phone check ──► evidence (JSON, verdicts already decided)
                                   │
                 for each gap ─────┤  prompt = evidence line + retrieved references + fix guidance
                                   ▼
                     fine-tuned model writes a JSON answer
                                   │
                     grounding check (groundingProblems)
                       ├─ passes ──► used in the report        (counted as "AI auditor")
                       └─ fails ───► built-in template instead (counted as "built-in writer")
```

The prompts are retrieval-style: everything the model may say (references, risk, fix steps, the found and required
values) is placed in the prompt, so the model phrases and selects rather than recalls. The same prompt builders exist in
Python (training) and JavaScript (app and phone page) and are tested to be identical (`app/test/prompt-parity.mjs`).

## The safety net

Every model answer is checked before it is used (`groundingProblems` in `app/renderer/auditor-core.js`, and
`training/check_faithfulness.py` for the training data):

- the answer is valid JSON with the required fields;
- every number in the answer appears in the evidence;
- the check id and status are the ones in the prompt;
- every reference (ISO control, GDPR or AI Act article) appears in the prompt;
- every AI tool named is one that was found.

If any rule fails, the built-in template answer is used for that item. The report states how many parts each wrote.
Because the verdicts come from code, **the audit is correct even without the model**; the model improves the
report's quality, not its correctness. This is deliberate: a compliance report must not depend on a model that can be
wrong.

## Current status

| | |
|---|---|
| Base model | Qwen2.5-0.5B-Instruct (Apache-2.0), quantised q4f16 / q4f32 for WebLLM |
| Base model without fine-tuning | 0% valid JSON, 0% grounded on 30 held-out examples: it wraps answers in Markdown, invents its own format and invents facts |
| Fine-tuned v0.1 pilot | **100% valid JSON, 100% grounded, 87% exact fields** on the same 30 examples (findings and AI classification 100%; summaries grounded but score / top gaps not exact) |
| Fine-tuning | LoRA on CPU (deliberately possible on an ordinary Mac). **In use now: run 1 on dataset v0.1** (pilot); next run on v0.3. See [Fine-tuning versions](#fine-tuning-versions) |
| Training data | `Devseis/endpoint-auditor-synthetic` (CC BY 4.0): synthetic devices, real CVEs from the vulnerability bundle, wording identical to the collectors and the phone page |
| Evaluation | `training/evaluate.py`: base vs fine-tuned on held-out data (JSON validity, grounding, status / reference accuracy) |

## Datasets and their versions

Two public datasets feed the model and the reports.

### Training data: [`Devseis/endpoint-auditor-synthetic`](https://huggingface.co/datasets/Devseis/endpoint-auditor-synthetic) (CC BY 4.0)

| Version | Where | Audits | Training examples | Covers | Used by |
|---|---|---|---|---|---|
| v0.1 | tag `v0.1` | 1,000 | 29,000 | Windows, Linux, macOS; 58 / 51 / 51 checks | **auditor-0.5b v0.1 pilot (training now)** |
| v0.2 | tag `v0.2` | 1,000 | 31,588 | adds signed-in user type, USB device access, pending updates, outdated apps, known vulnerabilities; 64 / 56 / 56 checks | not trained on (superseded) |
| **v0.3** | **main** | 1,500 | 39,500 | adds iPhone (29) and Android (30) checks, real CVEs, evidence labels, wording identical to the collectors and phone page | **auditor-0.5b v0.3 (next run)** |

Load a version with `load_dataset("Devseis/endpoint-auditor-synthetic", revision="v0.1")`. Local copies:
`training/out` (v0.1), `training/out-v0.2`, `training/out-v0.3`. Every example passes the faithfulness check
(`training/check_faithfulness.py`). The same seeds and the same vulnerability bundle give the same data.

### Vulnerability data: [`Devseis/endpoint-auditor-vulndb`](https://huggingface.co/datasets/Devseis/endpoint-auditor-vulndb) (CC BY 4.0)

| Version | Files | Used by |
|---|---|---|
| **2026-10-08** (current) | full bundle (4.9 MB: Linux advisories for Ubuntu, Debian, RHEL, AlmaLinux, Rocky; NVD ranges for 33 desktop apps and iOS / iPadOS; CISA KEV; EPSS) and a phone subset (55 KB, iOS / iPadOS); both Ed25519-signed | the desktop app, the phone page, and the v0.3 training data (real CVEs) |

This dataset is not tagged: each build replaces the previous one and is named by its build date, because apps should
always use the newest vulnerability data, and every report states the date it used. It goes stale as new
vulnerabilities are published, so it should be rebuilt about weekly (`python3 vulndb/build_bundle.py`, then
`node vulndb/sign-bundle.mjs`, then upload). Reports flag data older than 14 days. Training data records the bundle
date it was built with, so a dataset version stays reproducible.

## Fine-tuning versions

| Model version | Dataset | Training examples | Covers | Status |
|---|---|---|---|---|
| Base model (no fine-tuning) | — | — | — | Used by the app and phone page today; 0 of 30 answers grounded, so the built-in writer writes the report |
| **auditor-0.5b v0.1 (pilot)** | **v0.1** (tag `v0.1`) | 988 of 29,000 (balanced sample; 12 longer than 2,048 tokens skipped) | Windows, Linux, macOS; 58 / 51 / 51 checks | **Trained** 2026-10-08 (124 steps, final validation loss 0.0173). Evaluated: 100% valid, 100% grounded, 87% exact fields vs 0% for the base model. Published as `Devseis/endpoint-auditor-0.5b` tag `v0.1` |
| auditor-0.5b v0.3 | v0.3 (main) | 3,000 balanced, including phones (out of 39,500; none over 3,072 tokens) | Windows, Linux, macOS, iPhone, Android; 74 checks; real CVEs; evidence labels; computed summary scores | **Training** since 2026-10-08 21:20 (about 375 steps, about 17 hours). First version to ship |
| auditor-0.5b v0.4+ | v0.3 + expert-reviewed library + gold findings | more, on a GPU if available | same, better wording and priorities | Later |

Dataset v0.2 is not trained on: it was superseded by v0.3 before a run started.

**Settings (all runs):** Qwen2.5-0.5B-Instruct, LoRA r=16, alpha=32 (8.8 M trainable parameters, 1.75%), learning rate
2e-4, batch 1 with 8 gradient-accumulation steps, max 2,048 tokens, eager attention (sdpa gives wrong gradients with
PyTorch 2.2 on multi-threaded CPU), non-finite-gradient guard, checkpoints every 10 steps. About 2–2.5 minutes per step
(8 examples) on the 4-core Intel Mac.

### Next steps for fine-tuning

1. **Evaluate the pilot** (`training/evaluate.py`) on held-out v0.1 test data, base vs fine-tuned: valid JSON,
   grounding pass rate, correct status / check id / references. Decide whether 0.5B is enough.
2. **Train v0.3** from the base model (not on top of the pilot) on a balanced 3,000-example sample of v0.3 that includes
   iPhone and Android findings, summaries and AI classifications:
   `training/.venv/bin/python training/train_lora.py --data training/out-v0.3 --output training/models/auditor-lora-v0.3 --max-train 3000 --max-len 3072 --save-every 10 --eval-every 40`
   About 375 steps, roughly 17 hours on this Mac (resumable with `--resume`). A long limit is needed: 78% of Windows
   summaries (and 12% Linux, 5% macOS) are longer than 2,048 tokens. Before this run, the summary prompt was changed
   after the pilot evaluation: the pilot's summaries were grounded but its percentages were off by one and its top-gap
   lists only partly right, so the prompt now carries `SCORE (computed)` and `TOP GAPS (computed, highest risk first)`
   lines from code and the model quotes them instead of calculating (same change in Python and JavaScript; parity
   2,957 / 2,957). Python now rounds percentages half-up like JavaScript.
3. **Evaluate v0.3** on the v0.3 test split, per OS family, and on real runs (this Mac, the Ubuntu / Rocky containers,
   the phone page): the share of report items the model writes must be clearly above the pilot.
4. **Publish** if it passes: merge the LoRA adapter (`training/merge_lora.py`), convert to MLC (q4f16_1 and q4f32_1),
   upload `Devseis/endpoint-auditor-0.5b` (Apache-2.0) with a model card listing the dataset version and evaluation
   scores, and set `CUSTOM_MODEL` in `app/renderer/model-core.js` so the desktop app and phone page use it.
5. **Compare with 1.5B** on the same evaluation if 0.5B falls short (about 3× slower to train on CPU; heavier on phones,
   so 0.5B would stay the phone model).

Each published model version names the dataset version it was trained on, so results can be reproduced and cited.

## Running the model on devices (WebLLM)

| Build | Repository | Size | Used by | Measured (held-out prompts, WebLLM on an Intel Mac GPU) |
|---|---|---|---|---|
| q0f16 (unquantised) | [Devseis/endpoint-auditor-0.5b-q0f16-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q0f16-MLC) | 1.0 GB | desktop app, when the GPU supports 16-bit maths | v0.1: 6 of 6 grounded, 23–78 s per answer |
| q4f16_1 (4-bit) | [Devseis/endpoint-auditor-0.5b-q4f16_1-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q4f16_1-MLC) | 290 MB | phone check | v0.1: 5 of 6 grounded with the answer schema (0 of 6 without); full phone flow: 11 of 16 report parts written by the model (v0.1 never saw phone data) |
| q4f32_1 (4-bit) | [Devseis/endpoint-auditor-0.5b-q4f32_1-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q4f32_1-MLC) | 290 MB | GPUs without 16-bit maths | layout and format verified |

Each repository is tagged per model version (`v0.1`, `v0.3`); the app picks the version with `MODEL_VERSION` in
`app/renderer/model-core.js` and falls back to the base model if the download fails.

Lessons that shaped this:

- **Constrained decoding is required.** In 4-bit, the 0.5B model drifts on JSON syntax (stray brackets, a string where an
  object belongs), although the content is right. The app now passes a JSON schema per task (`ANSWER_SCHEMAS` in
  `app/renderer/auditor-core.js`) to WebLLM, so replies are always valid JSON in the expected shape. Without a schema,
  WebLLM 0.2.85's `json_object` mode fails to start at all ("Cannot pass non-string to std::string"), which had silently
  sent every desktop answer to the template writer.
- **Output length per task:** the AI-tool classification needs about 1,600 tokens; 900 cut it off.
- **Conversion:** MLC's own converter could not run here (mismatched Intel-Mac nightly builds; the stable Linux release
  needs an unpublished apache-tvm-ffi). `training/mlc_quantize.py` writes the MLC format directly and is checked against
  mlc-ai's builds of the base model: identical layout, q0f16 byte-identical, 4-bit scales bit-identical and 99.6-99.8% of
  4-bit values identical (the rest one step apart at rounding ties).

## How the model's role can be improved

Ordered roughly by value and effort. Each step keeps the rule that code decides verdicts unless stated otherwise, and
each is measured before it ships.

### 1. Better at what it already does (next)

| Improvement | How | Measure |
|---|---|---|
| Higher grounding rate | Train on v0.3; add hard cases (many gaps, long vulnerability lists, Error / Pending mixes) | Share of report items written by the model, on held-out data and on real runs |
| Better fix guidance | Expert review of `library/check_guidance.json` and `library/controls.json`, then retrain | Reviewer score on a sample of findings |
| More natural writing | Add paraphrase variety to the training templates; a small set of human-written gold findings | Human preference: model vs template |
| Prioritisation | Train the summary on vulnerability groups with KEV / EPSS so "fix first" reflects real exploitation | Agreement with a security reviewer's ordering |
| Size choice | Compare 0.5B with 1.5B on the same evaluation; keep 0.5B for phones if close | Grounding rate, speed, memory |

### 2. New roles that keep code as the judge

| Role | What the model does | Safeguard |
|---|---|---|
| **Q&A on the report** | Answers "why does this matter for GDPR?" or "how do I fix this on my Mac?" from the evidence and the library | Answers must cite the finding and library entries they use |
| **Phone interviewer** | Rephrases a question when the user is unsure, asks a follow-up, explains where to find a setting on their phone model | The answer recorded is still one of the defined options |
| **Audit comparison** | Explains what changed between two audits of the same device | Differences are computed by code; the model only describes them |
| **Policy drafting** | Drafts missing organisational documents (AI register entry, breach procedure outline) from the org answers | Marked as a draft for human review |
| **Translation** | Writes the report in the user's language | Same grounding rules; references stay untranslated |

### 3. Roles where the model judges evidence (later, with a second opinion)

| Role | What the model does | Safeguard |
|---|---|---|
| **Reading raw settings** | Interprets raw command output (an agent calling allow-listed read-only tools) instead of pre-digested values | The rule-based verdict still runs; disagreements are flagged for review, never silently resolved |
| **Screenshot reading on phones** | A vision model reads Settings screenshots instead of the user answering | Shown to the user for confirmation; labelled "read from screenshot" |
| **Unknown software** | Suggests which vulnerability product an unrecognised app is | Suggestion only; matching stays exact |

### 4. Learning from real use (with consent)

- Collect corrections from auditors on findings (opt-in, no device data leaves the organisation unless they choose to share
  anonymised findings) and add them to the training data.
- Track, per check, how often the model falls back to the template; retrain where it fails most.
- Publish each model version with its evaluation results on Hugging Face, so improvements are visible and citable.

## What the model will not do

- Decide a compliance verdict on its own.
- Invent CVEs, versions, numbers or references (the grounding check rejects them).
- Send anything about the device anywhere: it runs locally in the app or the browser.
- Replace an accredited auditor: the report supports ISO 27001 internal audits, GDPR Article 32 and EU AI Act deployer
  duties, but it is not a certification or legal advice.
