Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
Abstract
Deciphering animal intent is a fundamental challenge in computational ethology, heavily constrained by “semantic aliasing,” where identical external signals (e.g., a feline purr) map to vastly different internal states depending on physiological context. Existing Multimodal Large Language Models (MLLMs) are “modality blind” to high‑frequency biological time‑series data, restricting them to superficial behavioural pattern‑matching rather than genuine latent state reasoning. To bridge this gap, we introduce Meow‑Omni 1, the first open‑source quad‑modal MLLM purpose‑built for computational ethology. It natively fuses visual, audio, and physiological time‑series modalities with textual reasoning. Through targeted architectural model surgery, we integrate specialized scientific encoders into a unified backbone, formalizing intention inference via structural causal models. Evaluated on MeowBench, a novel expert‑verified quad‑modal benchmark, Meow‑Omni 1 achieves state‑of‑the‑art intent recognition accuracy (71.16 %), significantly outperforming leading vision‑language and omni‑modal baselines. We release the complete open‑source pipeline including model weights, training framework, and curated dataset to establish a scalable, robust paradigm for inter‑species communication and to advance foundation models toward real‑world veterinary diagnostics and wildlife conservation.