Unisound Launches Three Core Skills on ClawHub, Bringing Document Parsing, All-Scenario ASR and TTS to Power More Efficient AI Agents

Unisound 7
 Unisound Launches Three Core Skills on ClawHub, Bringing Document Parsing, All-Scenario ASR and TTS to Power More Efficient AI Agents

When documents can be understood in depth, speech recognition can accurately interpret context, and synthesized voices can convey warmth and a wide range of expression, entirely new possibilities emerge for AI agents in the workplace.

Effective immediately, three standardized core Skills from Unisound are now available in the ClawHub community. U2-doc-parser, U2-audio-file-transcriber, and U2-TTS bring high-accuracy document parsing, industry-leading all-scenario automatic speech recognition (ASR), and highly human-like text-to-speech (TTS) to the OpenClaw ecosystem. Together, they give developers the building blocks for AI agent workflows that can see clearly, listen closely, and communicate naturally - enabling agents to understand diverse documents, interpret varied speech, and respond appropriately to the context, while accelerating the transition of office automation and agent applications from prototype to industrial-scale deployment.

All three Skills draw on Unisound's core technology expertise in multimodal interaction and document intelligence and are powered by its proprietary large-model portfolio. They offer enterprise-grade reliability, rapid deployment, seamless orchestration, and deep adaptation to real-world scenarios. Developers can invoke them directly in OpenClaw without building services or managing environment dependencies, equipping agents with professional-grade 'eyes,' sensitive 'ears,' and an expressive 'voice' with minimal effort.

U2-doc-parser: High-Accuracy Document Parsing for Precise Information Extraction

Skill link: https://clawhub.ai/aaiccee/u2-doc-parser

As a standardized Skill powered by Unisound's U1-OCR large model, U2-doc-parser focuses on high-accuracy document parsing across a broad range of scenarios. It marks a leap from character perception to document understanding and serves as a core visual capability for agents handling complex documents, with particularly strong performance in professional contexts such as medical records, financial statements, and academic papers.

The Skill offers two core advantages:

1. Leading performance across multiple benchmarks: It delivers strong results in multiple authoritative evaluations, with clear advantages in demanding tasks such as table recognition, cross-page association, and small-text detection.

2. Semantics-driven structural understanding: Its pioneering 'semantic guidance + dynamic focus' strategy first analyzes the document structure and builds a 'semantic map,' much like a human expert. It accurately identifies the relationships among titles, charts, and body text. Even when faced with disordered layouts, mixed text and images, or multilingual content, it extracts information in a clear, organized manner, addressing the long-standing limitation of traditional OCR systems that can read words but cannot understand layout.

Within an agent workflow, U2-doc-parser can convert PDFs, images, and other document formats - including blurry photos, heavily watermarked pages, and bent or distorted documents - into structured Markdown data. The output can be consumed directly by downstream tasks without additional processing, making it well suited to office scenarios such as medical document processing, expense reimbursement review, and enterprise knowledge-base development.

U2-audio-file-transcriber: All-Scenario Speech Recognition - From Transcription to Understanding

Skill link: https://clawhub.ai/aaiccee/u2-audio-file-transcriber

Built on Unisound's Shanhai Zhiyin Large Model 2.0, U2-audio-file-transcriber serves as the agent's professional hearing core. It advances beyond basic speech transcription to deliver contextual understanding, domain adaptation, and compatibility across scenarios. The Skill can accurately recognize speech amid complex noise, dialectal accents, specialized terminology, and other challenging conditions - truly moving beyond hearing words to understanding meaning.

The Skill addresses key industry challenges through three core capabilities:

1. High accuracy in extreme conditions: In environments with complex background noise, recognition accuracy has exceeded 90% for the first time in the industry. Compared with mainstream ASR models, performance improves by 2.5%-3.6% in demanding scenarios involving complex noise and dialectal accents, supporting speech interaction across indoor near-field, noisy far-field, and public-space environments.

2. Broad multilingual and multi-dialect coverage: The Skill supports accurate transcription for more than 30 Chinese dialects and 14 international languages, from Cantonese, Hokkien, and Shanghainese to English, Japanese, Korean, and Thai. It also supports business meetings with mixed dialects and cross-border workplace communication.

3. Specialized terminology and contextual reasoning: Domain-specific terminology can be explicitly injected for targeted enhancement in healthcare, automotive, finance, and other fields. Examples include the drug names 'epalrestat' and 'metformin' in healthcare, and 'yoke steering wheel' in automotive applications, improving recognition accuracy by 30%. Powerful contextual reasoning also allows the Skill to infer critical information not explicitly stated, helping prevent breaks in meaning.

Within an agent workflow, U2-audio-file-transcriber can provide real-time transcription of meeting recordings, voice-command recognition in noisy environments, dialogue transcription in professional settings, and multilingual audio analysis. Transcription results can directly trigger subsequent agent actions, supporting workplace applications such as intelligent meeting assistants, voice-operated productivity tools, and automated customer-interaction records.

U2-TTS: Expressive Speech That Gives AI Warmth and Range

Skill link: https://clawhub.ai/aaiccee/u2-tts

Powered by Unisound's Shanhai Zhiyin Large Model 2.0, the intelligent U2-TTS acts as the agent's voice. With highly human-like delivery and creative versatility at its core, it combines realism with expressive flexibility, bringing greater warmth to technology and supporting workplace needs such as intelligent announcements, audio content creation, and scenario-specific voice interaction.

The Skill gives AI more expressive range through three core advantages:

1. Broad multilingual and multi-dialect coverage for context-appropriate expression: The Skill supports speech synthesis in multiple Chinese dialects and international languages. It produces authentic Cantonese and Sichuanese, while targeted optimization of features such as Japanese geminate consonants and Thai tonal variation delivers naturalness approaching that of native speakers. This makes it suitable for cultural and tourism promotion, cross-border workplace communication, dialect-based announcements, and more.

2. Diverse emotions and styles for authentic human expression: Users can switch among 12 Mandarin speaking styles, including gentle, capable, and friendly. The Skill can also naturally reproduce details such as laughter and breathing, and express emotions ranging from happiness and composure to urgency, allowing AI voice output to match the mood and atmosphere of different workplace scenarios.

3. Efficient creation and low-latency interaction across the workplace workflow: The Skill supports one-sentence voice cloning and can blend timbre and emotional characteristics from different voice samples to generate customized audio for narrated workplace content, video voiceovers, children's read-aloud content, and other uses. Its flow-matching module with causal-only attention and end-to-end fully streaming inference architecture significantly reduce system latency without compromising synthesis quality. In low-concurrency scenarios, time to first audio packet is reduced to under 90 milliseconds, delivering industry-leading real-time interaction for use cases such as intelligent voice announcements and real-time spoken responses.

Within an agent workflow, the intelligent U2-TTS converts an agent's text output into natural, context-appropriate speech. It can deliver spoken summaries of intelligent meeting minutes and voice reminders for reimbursement processes, extending agent interaction from text to speech and improving both efficiency and user experience in the workplace.

Rapid Integration and Seamless Orchestration for Industrial-Grade Agents

Unisound's three Skills were designed specifically for the OpenClaw ecosystem to minimize integration effort and cost, helping agent development progress from functional to truly effective:

1. Enterprise-grade reliability beyond the demo stage

All three Skills are built on Unisound's experience in real-world commercial deployments and have been validated at scale across healthcare, finance, office productivity, and other scenarios. Stable and predictable output, continuous official maintenance, and ongoing version upgrades help agents move beyond demos and into industrial-grade production environments.

2. Rapid deployment, ready out of the box

Developers can invoke the capabilities directly in OpenClaw as standardized Skill nodes, adding document parsing, speech recognition, and speech synthesis to an agent with one click - without substantial investment in technical R&D or environment setup.

3. Seamless composition and orchestration for customized intelligent workflows

The three Skills can be freely combined and flexibly orchestrated with other capabilities in the ClawHub ecosystem. As modular building blocks for agent development, they make it easy to create customized workplace agents:

Intelligent meeting assistant: U2-audio-file-transcriber transcribes meeting recordings and extracts key information, while U2-doc-parser analyzes meeting presentations, reports, and other documents. The agent automatically connects speech with document content to generate structured meeting minutes, which the intelligent U2-TTS can then present as audio. A task that once required an hour can be reduced to just a few minutes, enabling end-to-end intelligent meeting workflows.

Medical document processing agent: U2-doc-parser accurately analyzes medical invoices, expense lists, admission records, and other documents, extracts structured data, and applies business rules for compliance validation. The intelligent U2-TTS then announces the review results, enabling automated medical document processing with voice feedback.

Expense reimbursement agent: U2-doc-parser identifies and validates key information in reimbursement invoices and itemized lists, while U2-audio-file-transcriber captures an employee's spoken explanation. The agent automatically generates the reimbursement request, and the intelligent U2-TTS provides voice updates on its progress, creating a streamlined photo-plus-voice reimbursement process.

From high-accuracy document parsing and all-scenario voice interaction to human-like speech synthesis, Unisound's launch of these three core capabilities as standardized Skills on ClawHub marks an important deployment of its AI technology within the open-source ecosystem. It also provides a comprehensive foundation for office automation and agent development. Guided by the belief that 'true intelligence is not about showing off technology, but about integrating it into everyday life,' Unisound is working with the OpenClaw ecosystem to bring more efficient and intelligent AI agent applications to industries of every kind and redefine productivity in the intelligent workplace.

Starting today, developers can follow the Skill links on the ClawHub website and activate Unisound's high-quality document parsing, ASR, and TTS capabilities with one click - making it easy to build AI agent workflows that can see clearly, listen closely, and communicate naturally.