70 million conversations.
30 languages.

Our data captures nearly every global accent, covering a wide range of topics.

70M

CONVERSATIONS

8M+

HOURS

30

LANGUAGES

200+

COUNTRIES

ACCENT & DIALECT DIVERSITY

One language is never one voice

Real people don't speak in perfect, studio-recorded voices. Our data captures authentic regional accents and dialects from over 200 countries, giving you audio exactly how it sounds in the real world.

English

Neutral

Southern

Midwest

New England

Metropolitan

The West

AAVE

Pakistani

South African

Australian

Nigerian

Canadian

British RP

Filipino

Ghanaian

Singaporean

Scottish

New York

Kenyan

Irish

Caribbean

Indian

French

Metropolitan

West African

Central African

Maghrebi

Canadian

Belgian

Swiss

Caribbean

Spanish & Portuguese

Mexican

Caribbean

Central American

Colombian

Andean

Chilean

Rioplatense

Castilian

Canarian

US bilingual

Brazilian

European

Angolan

Mozambican

Arabic

Egyptian

Levantine

Gulf

Iraqi

Maghrebi

Sudanese

Modern Standard

German

Standard German

Austrian

Swiss

Northern

Bavarian-influenced

Asia

Mainland Mandarin

Taiwanese Mandarin

Cantonese-accented

Overseas communities

Delhi

Mumbai

Punjabi-influenced

Karachi

Lahori

Diaspora

French

Metropolitan

West African

Central African

Maghrebi

Canadian

Belgian

Swiss

Caribbean

Topic Distribution

Multi-Domain Conversational Coverage

Multi-Domain Conversational Coverage

Routing categories, topic labels, and human annotations for every conversation. The conversations run from international transfers to estate planning to firmware updates, each with its own vocabulary, stakes, and emotional register.

Money & Payments

International transfers

Payment failures

Refunds

Chargebacks

Fees & exchange rates

Billing disputes

Receipts & records

Identity & Security

Identity verification

Fraud reports

Account recovery

Suspicious activity

Compliance holds

Document review

Password resets

Legal & Business

Business formation

Trademarks & IP

Estate planning

Wills & trusts

Tax questions

Contracts

Annual filings

Registered agent

Devices & Technical

Troubleshooting

Setup & pairing

Firmware updates

App support

Compatibility

Repairs

Accessories

Orders & Logistics

Shipping & delivery

Returns & exchanges

Warranty claims

Order changes

Lost packages

Customs & international

Accounts & Relationships

Onboarding

Plan changes

Subscriptions & renewals

Cancellations

Escalations

Complaints

Retention & win-back

Profile updates

Technical Specs

Audio Format

• Dual-channel audio (agent + caller separated)
• 8kHz /16-bit PCM

Text Format

• Fully diarized with speaker-aligned timestamps
• Human-In-The-Loop golden transcripts (<2% WER)

Native Annotations
(Included)

Word-Level Timestamps

PII Redaction

Geography

CSAT

Intent

Sentiment

Overtalk

Human Annotations
(Available)

Emotions

Accent & Dialect

Code Switching

Disfluency

Social Dynamics

Cultural Patterns

Fully Compliant
Fully De-Identified

Fully Compliant
Fully De-Identified

All personally identifiable information has been removed from both audio and transcripts.
• The data is fully compliant with CCA, GDPR, and applicable privacy regulations
• Ready for use in model training, evaluation, and research without additional redaction or legal review.

Data
Origins

Data
Origins

Data Origins

Data Origins

• Originating from the global financial services industry, serving millions of customers across 200+ countries.
• Spanning account inquiries, transaction disputes, fraud resolution, compliance verification, and multilingual support interactions.

• Legally vetted data ownership and licensing rights

©️ 2026 Sumo AI

©️ 2026 Sumo AI

©️ 2026 Sumo AI

©️ 2026 Sumo AI