Multimodal AI
Multimodal AI is artificial intelligence that takes in more than one type of input — images, audio, text, documents — and reasons across them together. At AI USM this goes beyond emotion: the same system reads uploaded lab results, scans and records alongside the conversation itself. Emotion is one of several things it reasons about, and the multimodal emotion recognition page covers that part in detail.
Modalities AI USM works with
The platform accepts live camera and microphone input for emotion analysis, typed and spoken language, uploaded images, and documents such as health records and lab results. These are combined with what the memory layer already knows about the conversation.
- Vision: live camera analysis and uploaded images
- Audio: speech input and voice signal analysis
- Text: conversation in English, Russian, Chinese, Czech and Arabic
- Documents: PDFs and images of records, read for context
Why cross-modal reasoning matters
Reading a lab report and hearing worry in how a person asks about it are different inputs to the same question. Handling them in one system lets the answer address both: the numbers and the concern behind the question.
Limits
Multimodal does not mean unlimited. Quality varies by channel and by conditions, uploaded documents are read for context and not authenticated, and nothing in the pipeline produces a diagnosis.