logo
홈 뉴스

회사 소식 Scality's AI Inferencing Factory accelerates inferencing

인증
중국 Beijing Qianxing Jietong Technology Co., Ltd. 인증
중국 Beijing Qianxing Jietong Technology Co., Ltd. 인증
고객 검토
베이징 첸싱 지에통 테크 주식회사의 영업 사원은 매우 전문적이고 참을성 있습니다. 그들은 빨리 인용을 제공할 수 있습니다. 제품의 품질과 패키징은 또한 매우 좋습니다. 우리의 협력은 매우 매끄럽습니다.

—— 《Festfing DV》LLC

내가 긴급히 인텔 CPU와 토시바 SSD를 찾고 있었늘 때, 베이징 첸싱 지에통 기술 주식회사로부터의 샌디는 나에게 많은 도움을 주었고, 나에게 빨리 필요로 한 제품을 가져다 주었습니다. 나는 정말로 그녀를 압니다.

—— 고양이 엔

베이징 첸싱 지에통 기술 주식회사의 샌디는 내가 서버를 구입할 때 제시간에 나에게 구성 오류를 상기시킬 수 있는 매우 주의깊은 판매원을 있습니다. 엔지니어들은 또한 매우 전문적이고, 빠르게 테스팅 프로세스를 완료할 수 있습니다.

—— 스트렐킨 미하일 블라드미로비치

베이징 첸싱지에통과의 협업에 매우 만족합니다. 제품 품질이 훌륭하고, 배송도 항상 제 시간에 이루어집니다. 영업팀은 전문적이고, 인내심이 많으며, 모든 질문에 매우 친절하게 답변해 줍니다. 그들의 지원에 진심으로 감사드리며, 장기적인 파트너십을 기대합니다. 강력 추천합니다!

—— Ahmad Navid

품질: 제 공급업체와의 좋은 경험. 미크로틱 RB3011은 이미 사용되었지만 매우 좋은 상태로 모든 것이 완벽하게 작동합니다. 통신은 빠르고 원활했습니다.그리고 제 모든 걱정은 빠르게 해결되었습니다.매우 신뢰할 수 있는 공급자

—— 제란 콜레시오

제가 지금 온라인 채팅 해요
회사 뉴스
Scality's AI Inferencing Factory accelerates inferencing


Scality’s Maestro manages Artesca fleets. Its AI Inference Factory delivers a shared KV cache pool across on-premises GPU servers, matching Cloudian and MinIO while adding prefill/decode separation to boost GPU efficiency.


This AI Inference Factory combines validated open-weight models, a disaggregated inference-serving layer, and Scality AI Data Infrastructure (ADI). ADI uses policy-governed autonomous operations to manage large datasets across the full AI lifecycle. It gives enterprises, government agencies and neo-cloud providers a supported alternative to cloud AI services, removing the burden of assembling, integrating and maintaining the full software stack. The platform places ADI object storage under a vLLM/Dynamo serving layer as shared KV cache, with S3-over-RDMA and a native KV connector.


에 대한 최신 회사 뉴스 Scality's AI Inferencing Factory accelerates inferencing  0

Scality co-founder and CEO Jérôme Lecat said: “With AI moving into mission-critical production environments, organisations need greater control over where inference runs, how their models are managed and what happens to their data. For 15 years, Scality has built data infrastructure relied on by thousands of global customers for 24/7 operation. AI Inference Factory brings that experience to on-premises AI, giving organisations the reliability and sovereignty they need to run critical AI workloads on their own terms.”


The firm notes on-prem inference offers data sovereignty and potential cost benefits. Cloud inference uses per-token pricing, so costs rise as AI workflows grow more useful. Cloud providers may also alter model versions and quantisation at their discretion, disrupting dependent workflows.


Scality’s AI Inference Factory supports specialised chatbots, AI-assisted software development, agentic applications and other inference-heavy workloads. Its stack includes four software layers:
-Validated open-weight models maintained within the supported stack;
-Disaggregated inference serving separating prefill from decode so each can scale independently;
-A control plane that authenticates and meters requests, routes traffic to suitable GPUs holding relevant context and schedules workloads against SLA targets;
-Scality ADI, delivering high-performance object storage for models, enterprise data and inference state.


Scality launched Autonomous Data Infrastructure (ADI) in May as an enterprise-focused data infrastructure management platform. ADI uses policy-driven AI agent workers to place data across four storage tiers defined by performance, cost and protection.


ADI integrates with standard AI stacks and deploys at multi-petabyte scale on NVMe, with RDMA access to HDDs for nearly unlimited KV cache capacity. A single namespace automatically tiers data across TLC, HDD and optional tape, helping organisations match storage performance and cost to each stage of the AI data lifecycle.


ADI acts as the shared storage layer for AI Inference Factory, extending KV cache beyond GPU HBM limits. This lets AI inference infrastructure retain and retrieve model context rather than forcing GPUs to recompute it, lifting GPU utilisation and cutting inference costs.

Scality states ADI delivers a shared multi-petabyte cache readable by GPUs at latency within the same order of magnitude as GPU memory. Though slower, this allows GPUs to fetch existing context and focus on inference. This mechanism is similar to other KV caching systems.


에 대한 최신 회사 뉴스 Scality's AI Inferencing Factory accelerates inferencing  1


Co-founder and CTO Giorgio Regni said: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with GPUs staying fully occupied. Storage no longer forces KV cache to reside inside the GPU server.”


He added: “That enables disaggregated serving. One GPU pool handles prefill, processing prompts and writing KV cache to ADI. A second pool handles decode, reading cache back and generating tokens. With Scality ADI as shared cache, any decode GPU can pick up any context, with virtually no size limit. Prefill no longer interrupts decode, and both pools run at full load.”


As Regni describes, the AIIF architecture splits AI inference prefill and decode, using ADI as shared cache between GPU pools. Prefill computes KV cache for new prompts, while decode generates answers token by token. When both run on the same GPUs, new prefill jobs disrupt active decodes. Once separated, each pool handles one task full-time and scales independently, drawing from the same shared KV cache. Available decode GPUs can access stored context without binding it to a specific GPU server. Scality cites two examples showing performance gains from this separation:
DistServe (OSDI 2024) achieved up to 7.4 times more requests at identical latency by moving cache directly between GPUs. In Scality’s design, prefill writes KV cache to ADI and decode reads it back, so any decode GPU can access stored context, with tiered storage imposing almost no cache size limits.
Moonshot AI's Mooncake delivered 75 percent more requests on Kimi’s production traffic by deploying a shared KV cache pool.

Scality AIIF prefill-decode split.


Scality testing shows ADI can reduce GPU consumption while preserving performance:
1.9-second load time for Gemma-3 27B, nearly 10x faster than local NVMe, with the model streamed in parallel across the cluster over RDMA;
166 ms warm time-to-first-token restoring a 14K-token context from ADI, only 83 ms behind HBM;
14x faster KV cache retrieval than recomputation for a 14K-token context, up to 72x faster for a 439K-token context;
KV cache more than 80x larger than one GPU’s memory, supporting up to 1,000 resumable concurrent sessions without recomputation;
No measurable impact on token generation since context is restored before the first token;
97 percent of network line rate for transfers between GPUs and storage.


Scality validates and maintains the full stack. It works with open-source harnesses and agentic frameworks including OpenCode, Hermes, Goose, LangGraph and Pydantic AI, and runs validated open-weight models such as Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek. The software deploys on standard servers from Dell, HPE, Lenovo and Supermicro.


Stack components use open code, letting customers inspect how inference state is stored and moved and submit contributions for Scality review.


Scality AI Inference Factory's single end point.
Scality AI Inference Factory packages the whole serving stack behind one endpoint, including vLLM, Dynamo, the Scality AI Connector, Scality ADI and more. It is available now as a software license or fully managed service.


에 대한 최신 회사 뉴스 Scality's AI Inferencing Factory accelerates inferencing  2

                                                                           Nvidia KV caching


Nearly all filesystem storage vendors support Nvidia's STX scheme. It defines hardware requirements including BlueField-4 DPUs, Spectrum-X networking and dedicated flash storage tiers. Vendors in this group include DDN, Everpure, HPE, Hitachi Vantara, IBM, NetApp, Nutanix, VAST Data and WEKA. Object storage (S3) Nvidia CMX partners include Cloudian and MinIO with its petabyte-scale caching system.


Cloudian’s AIDP is a turnkey, on-premises appliance as an alternative to public cloud S3 data sources. This S3-compatible, RDMA-based datastore runs AI models and agents on Nvidia Blackwell GPU hardware and software plus BlueField DPUs, complying with Nvidia’s AI Data Platform reference architecture. It lets customers run production AI on their own IT infrastructure and retain direct control over sensitive data. Cloudian claims up to 60 percent cost savings by avoiding recurring token, egress and inference fees that make large-scale public cloud AI costly.


This value proposition resembles Scality’s, but lacks prefill-decode separation.
MinIO states MemKV lets an entire GPU cluster access a shared context pool at microsecond latencies matching inference speeds, rather than waiting for millisecond-scale latency. Again, this message aligns with Scality’s, without prefill-decode separation.


Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!

선술집 시간 : 2026-10-10 10:03:29 >> 뉴스 명부
연락처 세부 사항
Beijing Qianxing Jietong Technology Co., Ltd.

담당자: Ms. Sandy Yang

전화 번호: 13426366826

회사에 직접 문의 보내기 (0 / 3000)