Benchmark-Guided Load-Time Specialization for Local Language Models
SkillWeave is a technical proposal for loading local language models according to a measured workload and a device’s memory constraints. It favors a shared low-bit model base, selective precision and residency overlays, optional learned adapters, and rigorous benchmarks over duplicating a dense model for each specialist task.
Workload evidence can allocate a shared model’s limited precision and memory-residency budget to the parts that matter for a measured workload, while preserving the original model’s execution contract.
Copyright © 2026. The paper and summary are licensed under Creative Commons Attribution–NonCommercial–NoDerivatives 4.0 International. You may share unchanged copies with attribution for noncommercial purposes. Commercial use and adaptations require prior written permission.
This license covers the paper’s expression, not the underlying ideas, methods, or concepts. For implementation licensing, patent strategy, or commercial agreements, obtain qualified legal advice.