research/demo/02_eval_base.py
Builder 5dfc4a9fad Lead with a real end-to-end model-capability licensing demo
Rebuild the repo so its spine is a real, reproducible demonstration of
licensing an actual model capability, not payload-agnostic crypto on
stand-in blobs. The clean-room Ed25519 + AES-256-GCM primitives stay as
the fast mechanism layer; the real thing is now the headline.

New demo/ walkthrough (steps 1-7), each a standalone script printing
machine-checked evidence:
  1 download Qwen2.5-0.5B-Instruct from Hugging Face (gitignored cache)
  2 base scores 0.000 on an invented tool-call protocol (capability C)
  3 train a PEFT LoRA on C, base frozen (SHA-256 byte-identical proof)
  4 base + LoRA scores 0.925 on a held-out set with unseen arguments
  5 seal the adapter as an AES-256-GCM unit under an Ed25519 leaf cert
  6 valid licence decrypts-at-load and runs C at 0.925
  7 access-gate then crypto-erase: original and exfiltrated copy both
    permanently undecryptable, base alone back to 0.000

Reference run on an RTX 4090 captured the observed numbers now in the
README. keystore.py gains export_state/load_state so the authority (and
crypto-erasure) persists across the separate demo commands. A single
run_demo.sh drives steps 1-7; run_all.sh + pytest remain the fast
crypto-only mechanism tests.

Ships code only: base weights, HF cache, trained adapter, wrapping keys
and every sealed unit are gitignored and never committed. README rewritten
to lead with the demo and the observed numbers, with honest bounds
(in-memory adapter during a live licence needs a hardware enclave) and a
capability-tree scale-up as future work.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 03:20:12 +10:00

52 lines
1.8 KiB
Python

#!/usr/bin/env python3
"""Step 2: show the base model CANNOT perform capability C.
Runs the untouched base model on the held-out capability-C eval set and prints
its accuracy under the strict oracle. The prompt tells the model to emit a SIGIL
directive but reveals none of the protocol, so the base has no way to produce
the invented opcodes: accuracy is at or near zero. A few sample outputs are
printed so the failure is legible, not just a number.
"""
import json
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from demo.common import STATE_DIR, evaluate, load_base_model, load_tokenizer
from demo.dataset import build_dataset
def main() -> int:
print("=== Step 2: base model on capability C (expected: ~0) ===\n")
_, held_out = build_dataset()
print(f" held-out eval examples: {len(held_out)} "
"(locations and values disjoint from training)\n")
tokenizer = load_tokenizer()
model = load_base_model()
accuracy, correct, total, samples = evaluate(model, tokenizer, held_out, show=4)
print(" sample base outputs (request -> what the base produced):")
for request, gold, out, ok in samples:
print(f" request : {request}")
print(f" expected: {gold}")
print(f" base : {out!r} [{'ok' if ok else 'wrong'}]\n")
print(f" BASE ACCURACY ON C: {accuracy:.3f} ({correct}/{total})")
STATE_DIR.mkdir(parents=True, exist_ok=True)
(STATE_DIR / "eval_base.json").write_text(
json.dumps({"accuracy": accuracy, "correct": correct, "total": total}, indent=2)
)
if accuracy > 0.10:
print("\nRESULT: UNEXPECTED - base scored above 0.10; capability may not be novel.")
return 1
print("\nRESULT: PASS - the base model cannot perform capability C.")
return 0
if __name__ == "__main__":
sys.exit(main())