Shouya MaaS Intelligent Computing Platform

This enterprise‑LLM platform unifies heterogeneous computing, model assets, inference deployment and Token governance for controllable AI infrastructure.

Shouya MaaS Intelligent Computing Platform
Unified Resource Management
Unifies GPU/NPU heterogeneous resources via cluster‑server‑card hierarchy, centrally tracking resource specs, status and usage.
Standardized Model Deployment
Unifies model assets, links compute resources for deployment‑servitization base.
Inference Service Management
Deploys standardized inference services, manages instances, status and resources for stable enterprise‑model operation.
Fine-grained Token Statistics
Unifies model access, APIKey and Token usage, tracking requests, tokens, latency and exceptions for full invocation observability.

Industry Pain Points

After the enterprise AI moves from experimentation to large-scale application, computing power, models, services, and usage gradually become dispersed, making it difficult for traditional resource management methods to support unified operations.

It is difficult to unify heterogeneous resources

GPUs and NPUs are scattered across different clusters and servers, lacking a unified view of resource specifications, operational status, and utilization, leading to a continuous increase in management costs.

Model assets are difficult to manage

The maintenance of weight files, model images, and runtime configurations is fragmented, and there is a lack of unified standards for model versions and computing power requirements, making deployment preparation complex.

The efficiency of model deployment is low

Model deployment relies on manual judgment of hardware specifications, video memory, and operating environment, lacking a standardized basis for matching resources with models.

Model usage is difficult to manage

After continuous invocations from multiple models and applications, the request volume, token consumption, and API key usage become dispersed, making it difficult to uniformly track costs and abnormal invocations.

Core Product Matrix

Covering computing power resources, model assets, inference services, and model invocation, we establish a complete management chain for enterprise large models, spanning from resource access to service operation.

01

Computing Power Resource Management

One-stop heterogeneous computing management. Adopt auto-discovery, manual & batch import for servers. Build mapping among clusters, servers and cards. Standardize runtime specs via templates to support full lifecycle computing scheduling.

Resource Overview: Display cluster, server & card quantity and utilization
Server Management: Support auto-discovery, manual and batch import
Card Management: Identify GPU/NPU; monitor model, temperature, power and deployment relations
Computing Template: Standardize hardware, software and network requirements for pre-deployment verification
Computing Power Resource Management
02

Model weight file

Build a deployable asset system that spans from model files, images, to versions. The platform registers model files through server paths and automatically identifies their attributes, and registers model images according to specifications. The combination of file and image registration generates models and versions. This module provides standardized and reusable model versions for inference services, ensuring that each deployment has clear source, environment, and version baselines.

Model weight file: Supports both server synchronization and file upload methods, for unified management of existing model weight assets in the enterprise
Model image management: Unify the registration and management of model running images, providing a standard operating environment for model deployment.
Model Release: Bind files and images, configure runtime specs to create model versions
Version Management: Support multi-version iteration; deploy versions directly to inference services
Model weight file
03

Resource Matching

Quickly launch running services with selected models. Four-step guided deployment with seven pre-checks. Decouple services and instances, support elastic scaling and real-time monitoring.

4-step Deployment Guide: Model selection → resource pool → 7 checks → service deployment
Resource Validation: Match runtime specs, output compatibility results and alerts
Real-time Monitoring: Track instance status, resource usage and trigger anomaly alerts
Resource matching: Filtering computing resources that meet the requirements of card type, video memory, and card quantity based on the computing power template associated with the model
Resource Matching
04

Token Hub

Unify AI service gateway, full lifecycle API Key management and Token statistics. Form closed-loop call tracking without parsing request body, ensure access security and compliance.

Access Gateway: Unified service entry, configure address, authentication and identity mapping
API Key Management: Full lifecycle control, manage access scope and Token quota
Token Statistics: Multi-dimensional analysis of call volume, consumption, model & user ranking
Token Details: Request-level logs, filter and audit by API Key, model and time
Token Hub

Core Advantages

Visible computing power, controllable deployment, traceable consumption

Full-link Bidirectional Tracing

Bidirectional tracing via 5D relations: trace resources from services and vice versa.

Mandatory Pre-deployment Check

Standardize runtime specs; 7 pre-deployment checks eliminate resource mismatches.

Decouple Services & Instances

Separate services and instances; support elastic scaling with auto resource verification.

Decouple Statistics & Inference

Token Hub independently tracks requests and Token consumption, supports multi-dimensional aggregation.

Build Full-lifecycle Management Platform for Enterprise Private Large Model Services

Consult Shouya MaaS now, contact our experts for further support.

Contact Us