Hetzner已开始提供一项实验性的大语言模型推理服务,名为Hetzner Inference [1]。该服务采用OpenAI兼容的API接口,运行在Hetzner自有基础设施之上 [1]。
目前该服务仅搭载一个可用模型——Qwen/Qwen3.6-35B-A3B-FP8,拥有35B参数和3B活跃参数 [1]。该模型接受文本和图像输入,具备262K的上下文窗口,采用FP8量化权重 [1]。根据2026年7月23日的测试结果,该模型的中位首token延迟为153毫秒,输出吞吐量为224 token/秒(设置512 token上限)[1]。
Hetzner当前公开的GPU服务器配置包括RTX 4000 SFF Ada(20GB)和RTX PRO 6000 Blackwell Max-Q(96GB)[1]。
该服务目前处于早期实验阶段,不提供计费、不提供服务级别协议保障,也不提供生产环境保证 [1]。Hetzner通过此举对用户需求、系统可扩展性和负载处理能力进行测试 [1]。
Hetzner has unveiled Hetzner Inference, an experimental large language model inference service built on its own infrastructure and offering an OpenAI-compatible API [1]. The platform currently runs the Qwen3.6-35B model, which features 35 billion parameters with 3 billion active parameters and accepts both text and image inputs [1].
The service is in early-stage testing with no billing charges, service-level agreements, or production guarantees [1]. Through this offering, Hetzner is assessing user demand, system scalability, and load-handling capabilities [1].
Performance benchmarks from testing conducted on July 23, 2026, show a median time-to-first-token of 153 milliseconds and output throughput of 224 tokens per second, with a 512-token output limit [1]. The Qwen model operates with FP8 quantized weights and provides a context window of 262,000 tokens [1].
Hetzner's publicly available GPU server configurations include the RTX 4000 SFF Ada with 20GB memory and the RTX PRO 6000 Blackwell Max-Q with 96GB memory [1].