AI Agent Infra System
Integrated edge deployment with AetherStation
01/ HARDWARE & INFRA
AetherStation

SOFTWARE
EDGE AI AGENT INFRA
HARDWARE
Editions VERSION
A / Domestic Edition (Xinchuang)
AetherStation Domestic Edition
A domestic, Xinchuang-ready agent appliance
Full-stack domestic chips · Data stays on premises Government / Finance / Healthcare
- Core Configuration
- Moore Threads GPU + Houmo NPU ×2 370 TOPS INT8 of domestic heterogeneous compute
- Typical Use Cases
- Government documents · Financial compliance review Medical record analysis · Manufacturing knowledge bases
B / Standard Edition (NVIDIA)
AetherStation Standard
High-performance agent workstation
Full-Stack Domestic Chip Support · Data Stays On-Premises Government / Finance / Healthcare
- Core Configuration
- NVIDIA GPU + AetherInfer Engine Qwen3.6-27B: 120 tok/s single-stream; 630 tok/s at 16-way concurrency 35B MoE · 20 concurrent users (for teams of up to 50)
- Typical Use Cases
- Coding agents · Data analysis Enterprise workflows · Enterprise assistants
02 / ENTERPRISE WORKFLOWS
ATTUNE: Enabling Small Local Models to Handle Complex Enterprise Workflow Tasks
Without modifying the base model's weights, Attune achieves Teacher-level performance across 15 workflow tasks, with perfect held-out scores on 13 of them.
Workflow Task Pass Rate (%)
- Compliance Alert
- Transcript Extraction
- Smart Home
- In-Vehicle System Control
- Government Service Routing
- Customer Service Routing
- IT Operations Work Orders
- Contract Data Extraction
- Medical Record Extraction
- Industrial Edge
- Quality Inspection Classification
- Credit Data Extraction
- Claims Data Extraction
- Expense Data Extraction
- Robot Commands
03 / INFERENCE PERFORMANCE
AetherInfer Engine: Unlock Production-grade Inference Performance
Designed for high-throughput, low-latency, highly concurrent agent workloads, AetherInfer co-optimizes hardware and AetherHeart’s proprietary inference stack to deliver more real-world performance per unit of hardware investment.
- 12.58×︎
- Single-batch Output Speed 124 tok/s vs 9.86 tok/s124 vs 9.86 tok/s
- 4.92×︎
- Total Throughput at 16-way Concurrency 630 tok/s vs 128 tok/s630 vs 128 tok/s
- 210 16 tok/s per ¥︎10K16 tok/s / ¥︎10K
AdvantageCo-optimization of hardware and our proprietary inference stack—not simply adding more GPUs
04 / THEORETICAL EFFECT
For common enterprise tasks, the cost per task does not decrease linearly.
100 active employees · Approximately 2 million workflows annually · All delivering Attune-verified equivalent task outcomes
1.0XAetherStation Benchmark
| Deployment Options | Cost per Completed Task | Ops Overhead | Upfront Cost |
|---|---|---|---|
| 8×A100 · Qwen 27B | 30–40× | Medium | ¥900,000–¥1,200,000 |
| 8×H100 · Qwen 27B | 11.9× | High | ¥2,200,000–¥3,000,000 |
| GLM-5.2 · Official API | 11.7× | Low | API Usage Fees |
| 12× B300 · Officially Optimized by NVIDIA · GLM-5.2 | 10.3× | High | ¥10M+ |
| 6×AetherStation · Qwen 27B | 1.0× | Low | ¥180,000 |
Estimation note: The figures above are a theoretical comparison based on assumptions regarding task volume, model, concurrency, hardware, and service pricing. They are intended to illustrate the cost structure; actual results may vary depending on deployment conditions.
