AI Agent Infra System

Integrated edge deployment with AetherStation

01/ HARDWARE & INFRA

AetherStation

AetherStation 边缘智能体工作站

SOFTWARE

TamerDynamic task understanding, runtime orchestration, and execution state machines

EDGE AI AGENT INFRA

AttuneAgent capability tuning and edge deployment hub
AetherInfer EngineQuantization / Speculative Decoding / KV Cache Management

HARDWARE

ChipsHeterogeneous compute across NVIDIA, Ascend, Moore Threads, Intel, and more

Editions VERSION

A / Domestic Edition (Xinchuang)

AetherStation Domestic Edition

A domestic, Xinchuang-ready agent appliance

Full-stack domestic chips · Data stays on premises Government / Finance / Healthcare

Core Configuration
Moore Threads GPU + Houmo NPU ×2 370 TOPS INT8 of domestic heterogeneous compute
Typical Use Cases
Government documents · Financial compliance review Medical record analysis · Manufacturing knowledge bases

B / Standard Edition (NVIDIA)

AetherStation Standard

High-performance agent workstation

Full-Stack Domestic Chip Support · Data Stays On-Premises Government / Finance / Healthcare

Core Configuration
NVIDIA GPU + AetherInfer Engine Qwen3.6-27B: 120 tok/s single-stream; 630 tok/s at 16-way concurrency 35B MoE · 20 concurrent users (for teams of up to 50)
Typical Use Cases
Coding agents · Data analysis Enterprise workflows · Enterprise assistants

02 / ENTERPRISE WORKFLOWS

ATTUNE: Enabling Small Local Models to Handle Complex Enterprise Workflow Tasks

Without modifying the base model's weights, Attune achieves Teacher-level performance across 15 workflow tasks, with perfect held-out scores on 13 of them.

15 Industry Tasks, 13 Perfect Scores

Workflow Task Pass Rate (%)

  1. Compliance Alert
  2. Transcript Extraction
  3. Smart Home
  4. In-Vehicle System Control
  5. Government Service Routing
  6. Customer Service Routing
  7. IT Operations Work Orders
  8. Contract Data Extraction
  9. Medical Record Extraction
  10. Industrial Edge
  11. Quality Inspection Classification
  12. Credit Data Extraction
  13. Claims Data Extraction
  14. Expense Data Extraction
  15. Robot Commands

03 / INFERENCE PERFORMANCE

AetherInfer Engine: Unlock Production-grade Inference Performance

Designed for high-throughput, low-latency, highly concurrent agent workloads, AetherInfer co-optimizes hardware and AetherHeart’s proprietary inference stack to deliver more real-world performance per unit of hardware investment.

12.58×︎
Single-batch Output Speed
124 tok/s vs 9.86 tok/s124 vs 9.86 tok/s
4.92×︎
Total Throughput at 16-way Concurrency
630 tok/s vs 128 tok/s630 vs 128 tok/s
210
16 tok/s per ¥︎10K16 tok/s / ¥︎10K
Throughput per ¥10K InvestedDGX ¥︎31,900 · Aether ¥︎30,000
AetherStation Standard
210
AetherStation Lite
110
DGX Spark · MTP
72.6
DGX Spark · no MTP
40.1

AdvantageCo-optimization of hardware and our proprietary inference stack—not simply adding more GPUs

04 / THEORETICAL EFFECT

For common enterprise tasks, the cost per task does not decrease linearly.

100 active employees · Approximately 2 million workflows annually · All delivering Attune-verified equivalent task outcomes

1.0XAetherStation Benchmark

Comparison of per-task cost, operational complexity, and upfront cost across enterprise deployment options
Deployment OptionsCost per Completed TaskOps OverheadUpfront Cost
8×A100 · Qwen 27B30–40×Medium¥900,000–¥1,200,000
8×H100 · Qwen 27B11.9×High¥2,200,000–¥3,000,000
GLM-5.2 · Official API11.7×LowAPI Usage Fees
12× B300 · Officially Optimized by NVIDIA · GLM-5.210.3×High¥10M+

Estimation note: The figures above are a theoretical comparison based on assumptions regarding task volume, model, concurrency, hardware, and service pricing. They are intended to illustrate the cost structure; actual results may vary depending on deployment conditions.