Fermion Research 推出了 Neutrino-1 8B,一款参数量为 8.19B 的开源大语言模型 1。该模型采用专有的三进制权重格式,将文件大小压缩至 3.88 GB 1,可在数据中心 GPU、Mac 和桌面 CPU 上运行 1。
模型在不同硬件上的推理性能表现各异 1。在 Apple M5 CPU 上,单流推理速度达到 24.9 tokens/秒(9 线程配置);NVIDIA L4 GPU 可实现 30.7 tokens/秒;Apple Silicon MacBook 的推理速度最快,达到 33.7 tokens/秒 1。在投机式解码任务中,0.6B 草稿模型在计数提示上的接受率达到 100%,在事实提示上的接受率为 96.5% 1。
该模型衍生自 Qwen3-8B 1,已在 Apache License 2.0 下开源发布,允许商业使用、修改、微调和再分发 1。
Fermion Research has unveiled Neutrino-1 8B, an open-source large language model with 8.19 billion parameters, featuring a proprietary ternary weight format that compresses the model file size to 3.88 GB 1. The compact model enables deployment across diverse hardware platforms, from data center GPUs to Mac systems and desktop CPUs 1.
The model demonstrates competitive inference performance across multiple configurations 1. On Apple M5 CPUs with nine threads, Neutrino-1 8B achieves 24.9 tokens per second, while NVIDIA L4 GPUs deliver 30.7 tokens per second with a 4.68 GB context window 1. Apple Silicon MacBooks reach 33.7 tokens per second 1. In speculative decoding benchmarks, a 0.6B draft model shows a 100% acceptance rate on counting prompts and 96.5% on factual prompts 1.
The model is released under the Apache License 2.0, permitting commercial use, modification, fine-tuning, and redistribution 1. Neutrino-1 8B is derived from Qwen3-8B, which similarly uses the Apache-2.0 license 1.
评论
还没有评论,欢迎留下第一条。