Devseis Endpoint Auditor — stage 1, steps 1–3

A read-only audit of a Windows, Linux or macOS endpoint against a security baseline, GDPR technical controls, EU AI Act controls and organisational controls. It produces:

This repository holds steps 1–3 of the MVP: the collectors, the evidence format, the check catalog, a synthetic evidence generator, the control library and the training data builder, so the model can be built before real customer data exists. The generated training data is in the dataset Devseis/endpoint-auditor-synthetic.

Not a certification. ISO 27001 certification needs an accredited auditor, and much of GDPR and AI Act compliance is organisational. The tool collects evidence and flags gaps.

What it checks

Section Checks Examples
COMPLIANCE RESULTS Security baseline (33 on Windows, 26 on Linux/macOS) Password and lockout policy, antivirus, firewall, screen lock, USB storage, administrators, disk encryption
GDPR TECHNICAL CONTROLS 12 Automatic updates, supported OS, security logging, time sync, backup, removable media encryption, guest account, remote access, diagnostic data, personal data discovery (opt-in, later stage)
EU AI ACT CONTROLS 5 AI software inventory, unapproved AI tools (shadow AI), risk classification (by the model), Windows Recall / Apple Intelligence, AI log retention
ORGANISATIONAL CONTROLS 8 RoPA, DPIA, breach procedure, processor agreements, AI register, AI literacy, human oversight, AI transparency

Every check, with its ISO 27001:2022 Annex A, GDPR and AI Act references, is in catalog/checks.json. Control titles are referenced; the ISO standard text is not reproduced (it is copyrighted).

Statuses: Compliant, Non-Compliant, NotApplicable, Pending (needs the model or an organisational answer) and Error (value not readable, usually because the collector was not run as administrator/root).

Repository layout

baseline.conf                     required values (password length, lockout, USB, approved admins, AI tools…)
organisational.conf               answers to the organisational questions (yes / no / partial / n/a)
catalog/checks.json               every check: id, section, name per OS, rule, control references
schema/evidence.schema.json       JSON Schema of the evidence file
collectors/windows/audit.ps1      Windows collector (Windows PowerShell 5.1 or PowerShell 7)
collectors/linux/audit.sh         Linux collector (Ubuntu/Debian, RHEL family)
collectors/macos/audit.sh         macOS collector (bash 3.2 as shipped with macOS)
synthetic/generate_evidence.py    synthetic evidence in the same format, using the same rules
synthetic/samples/                three example synthetic records (one per OS)
tools/validate_evidence.py        checks evidence against the schema rules and the catalog
tools/render_report.py            turns evidence back into the text report
library/controls.json             ISO 27001 controls, GDPR and AI Act articles in Devseis's own words
library/check_guidance.json       per check: why it matters, risk, fix steps per OS, how to verify
library/ai_tools.json             AI tools the inventory can find: data location, default AI Act class, GDPR notes
training/build_dataset.py         evidence -> chat-format training examples (finding, summary, AI classification)
training/check_faithfulness.py    checks that every answer is grounded in its prompt

Running an audit

The collectors only read settings. The only files they write are the report and evidence (plus short-lived exports from secedit/auditpol in TEMP on Windows, deleted straight away).

# Windows (elevated PowerShell). Output: C:\ProgramData\Devseis\EndpointAudit\
powershell -ExecutionPolicy Bypass -File .\collectors\windows\audit.ps1
# Linux. Output: /var/log/devseis-endpoint-audit/
sudo ./collectors/linux/audit.sh

# macOS. Output: /Library/Logs/Devseis/EndpointAudit/
sudo ./collectors/macos/audit.sh

Options: --baseline FILE, --org FILE, --output DIR (Windows: -Baseline, -Organisational, -Output). Exit code: 0 all compliant, 1 at least one Non-Compliant, 2 could not run or write output.

Without administrator/root rights the audit still runs; checks that need those rights report Error.

Configuring the baseline

Edit baseline.conf. Important values to set per organisation:

Answer the eight questions in organisational.conf once; they appear in every report.

Synthetic data

python3 synthetic/generate_evidence.py --count 600 --seed 7 --out synthetic/out/evidence.jsonl
python3 tools/validate_evidence.py synthetic/out/evidence.jsonl
python3 tools/render_report.py synthetic/out/evidence.jsonl --index 0

Records follow the collectors' check order, names and pass/fail rules, across three device profiles (hardened, typical, neglected), current and unsupported OS versions, runs without admin rights, third-party antivirus, and AI tool inventories. Host names, serial numbers and accounts are invented. The same seed always gives the same data.

Control library and training data (steps 2–3)

The library is what the model retrieves at run time; the training examples teach it to turn evidence plus library entries into findings. Three tasks are built from every audit record:

Task Input Output (JSON)
finding one check: status, found, required, references, guidance title, finding, risk, OS-specific fix steps, verification, references
summary all results of one audit score, counts, top 5 gaps by risk, summary, next steps
ai_classification AI tools found, approved list, tool facts, AI Act rules per tool: default risk class, when it becomes high-risk, GDPR points, action
python3 synthetic/generate_evidence.py --count 1000 --seed 7 --out synthetic/out/evidence.jsonl
python3 training/build_dataset.py --evidence synthetic/out/evidence.jsonl --out training/out
python3 training/check_faithfulness.py training/out/*.jsonl

Labels come from templates, not from a language model, so they never contain invented values. Splits are made per audit (80/10/10), so no device appears in two splits. With 1,000 synthetic audits this gives 29,000 examples (23,083 train / 3,123 validation / 2,794 test); all pass the faithfulness check.

Template labels teach the format and grounding. Before customer use, add reviewed real audits and have an ISO 27001 / GDPR professional review the library and a sample of the targets.

Desktop app (Electron + WebLLM)

app/ is the client application. It runs everything on the client's computer:

  1. Prepare — downloads the auditor model once through WebLLM (WebGPU) and caches it for offline use. Without WebGPU, the built-in report writer is used instead.
  2. Approve — lists every check that will run for this OS; the client clicks Run audit and the OS asks for administrator rights (macOS admin prompt, Windows UAC, Linux pkexec).
  3. Progress — each check is shown as it finishes (collectors write --progress lines), then the findings are written.
  4. Save — a popup asks where to save the PDF report; with no answer in 60 seconds it is saved to Documents/Devseis Endpoint Audit/<computer>-<date>/ together with the evidence JSON.

Model answers are only used if they pass the same grounding check as the training data; otherwise the template answer is used, so a report is always produced. Scores and counts always come from the evidence.

cd app && npm install && npm test   # prompt parity with training/build_dataset.py
npm start                           # run the app

Offline fine-tuning on CPU: training/train_lora.py (LoRA on Qwen2.5-0.5B-Instruct; see --benchmark and --resume).

Role of the language model

Code collects the evidence and decides every verdict; the fine-tuned model explains, prioritises and writes the report, and every answer is checked against the evidence before it is used. See docs/MODEL_LLM.md for who does what and how the model's role can grow.

Phone check (iPhone, iPad, Android)

The Space's front page (index.html) is a phone self-check that runs entirely in the browser:

  1. Detects what a browser can see: OS and version (Safari 26+ hides the iOS version in its user agent, so Safari's own version is used; Android uses client hints), device model (Android), browser, whether a device unlock method is set up (WebAuthn), and WebGPU support.
  2. Asks the rest, one question at a time with where to look in Settings (catalog/mobile_questions.json). Every result is labelled detected or self-reported in the evidence and the report.
  3. Matches the iOS version against the signed phone subset of the vulnerability bundle (55 KB, signature checked with WebCrypto).
  4. Writes the findings with the same auditor model in WebLLM (about 350 MB download, roughly 1 GB of GPU memory, so newer phones) or with the built-in writer, and builds the same report as the desktop app.
  5. Saves or shares the report as a PDF (share sheet or download). Phones do not allow a silent auto-save.

mobile/mobile-core.js turns detections and answers into evidence; the synthetic generator calls the same code through mobile/build-evidence.mjs, so training data has exactly the page's wording. Organisation answers can be passed in the link: index.html#org=<base64 JSON {"ORG_ROPA": "yes", ...}>.

Test without a phone: serve the repo (python3 -m http.server 8765) and run app/node_modules/.bin/electron mobile/test/e2e.cjs ios out.json (or android); it answers every question in a phone-sized window with a phone user agent and builds the PDF; then check the evidence with python3 tools/validate_evidence.py out.json.

Vulnerability matching (offline)

vuln.known_vulnerabilities is decided by exact version comparison, never by the model:

  1. The collector saves a software inventory next to the evidence (*-inventory.json: names and versions only; it stays on the computer).
  2. The app matches it against a signed vulnerability bundle (app/vulndb.cjs): Linux packages against OSV advisories for Ubuntu, Debian, RHEL, AlmaLinux and Rocky (with dpkg/rpm version rules), Windows and macOS applications against NVD version ranges for 33 common business apps (vulndb/products.json), with CISA KEV ("actively exploited") and EPSS scores for priority.
  3. Only vulnerabilities with an available fix are counted. Distributions or apps outside the bundle are reported as not covered, never as "0 vulnerabilities". Data older than 14 days is flagged in the finding.

Devseis builds the bundle with python3 vulndb/build_bundle.py (needs internet; about 1 GB of source downloads, 5 MB result) and signs it with node vulndb/sign-bundle.mjs (private key in ~/.devseis/, never in the repository). Test a run without the app: node tools/match_vulns.cjs <evidence.json>. Packaging must ship vulndb/out/ inside the app's audit resources.

Version status

Part Version Notes
Catalog, library, synthetic data, training data 0.3 74 checks: Windows 64, Linux 56, macOS 56, iPhone 29, Android 30. 0.3 adds phones, real CVEs in training data and wording identical to the collectors and the phone page; 39,500 grounded examples
Collectors (Windows, Linux, macOS) 0.2 (catalog 0.3) macOS tested on a real Mac; Linux tested on Ubuntu 24.04 and Fedora 42 containers; Windows parse-checked only (needs a real Windows test). Known vulnerabilities: inventory only, matching comes with the offline vulnerability bundle
Vulnerability matcher and bundle 0.1 Tested: this Mac (Chrome 154 → 58 critical/high, fixed in 155), Ubuntu 24.04 (OpenSSL), Rocky 9 (25 of 25 packages agree with dnf's own security list), Fedora reported as not covered
Phone check (web) 0.3 Tested in a phone-sized window as iPhone (Safari 26) and Android (Chrome 141); needs testing on real phones
Fine-tuned model v0.1 published, v0.3 training Devseis/endpoint-auditor-0.5b (+ q0f16 / q4f16_1 / q4f32_1 WebLLM builds), tag v0.1: 100% valid and grounded on held-out data vs 0% for the base model; used by the app and the phone check. v0.3 (phones, computed summary scores) is training and replaces it when evaluated

Test status (collector version 0.1.0)

Collector How it was tested
macOS Run on a real Mac (macOS 26, Intel) without root: evidence valid; report identical to the rendered evidence apart from the time-zone label.
Linux Run in an Ubuntu 24.04 container as root and as a normal user: evidence valid, results checked by hand. Still to test on a full desktop install (GNOME, LUKS, ufw).
Windows Parsed with PowerShell 7 (0 errors) and checked with PSScriptAnalyzer (no errors). Still to run on a real Windows 11 machine.
Synthetic 600 records, all valid against the schema and catalog; reproducible from the seed.

Privacy

Real evidence contains the host name, serial number, user account names and file paths: it is personal data under GDPR. Keep it on the device or in the organisation's own storage. Only synthetic data belongs in this repository.

Roadmap

  1. Collectors, evidence format, catalog, synthetic data ← done
  2. Control library in Devseis's own wording, for retrieval by the model ← done (expert review pending)
  3. Training data: evidence → auditor findings (explanation, risk, recommendation, references) ← done on synthetic data; expert review of a sample pending
  4. Fine-tune a small open-weight model (LoRA/QLoRA) to run locally
  5. Evaluate on held-out data (correct references, no invented values) and publish the scores
  6. Package: one installer per OS that detects the OS, runs the collector and writes the report locally

Licence and citation

Open source by Devseis:

Part Licence
Code (collectors, app, training, tools, vulndb builder) Apache 2.0
Data and content (catalog, control library, fix guidance, baseline, docs, synthetic and training data, vulnerability bundle) CC BY 4.0

Keep the NOTICE file when redistributing, and credit "Devseis Endpoint Auditor by Devseis". To cite it, use CITATION.cff:

@software{devseis_endpoint_auditor_2026,
  author  = {{Devseis}},
  title   = {Devseis Endpoint Auditor: a local, read-only endpoint auditor for ISO 27001, GDPR and the EU AI Act},
  year    = {2026},
  version = {0.2.0},
  url     = {https://huggingface.co/spaces/Devseis/endpoint-auditor},
  license = {Apache-2.0}
}

Contributions are welcome: open a discussion on the Space. Reviews of the control library and fix guidance by security practitioners are especially useful.