Hello. Ask me about angular attention, ternary weights, retrieval, or the limits of this proposal.
Talk to CRATER Tiny
CRATER Tiny
Ask about angular attention, ternary weights, retrieval, or the proposal's limits.
Simulation, not a trained checkpoint. Responses come from a small local knowledge map. The operator trace shows data flow, not hidden chain-of-thought. No network request occurs.
Proposed LLM artifact
Specification preview
crater-tiny.crater
No model file is loaded by this demo. A deployed runtime would verify the signed manifest once, memory-map the packed weights, then reuse the loaded artifact for every message.
Open file→Verify manifest→mmap weights→Run tokens
- Packed core
- Base-3 ternary weights plus 2-bit quaternary heads Q/K/V/O + experts
- Quantization metadata
- FP16 output scales γw, learned thresholds Δℓ, head mask Applied after accumulation
- Sensitive modules
- Embeddings, CLN, router, gates, LoRA and output head FP16 / optional INT8
- Manifest + tokenizer
- Architecture, packing layout, IRW settings and vocabulary Signed configuration
The paper's 7B-complex example estimates about 2.8 GB for packed ternary weights and about 3 GB after stated overhead. CRATER Tiny has no released checkpoint or measured file size.