Research &
Publications.
Peer-reviewed survey papers, journal articles, and investigative studies in speech transformers, large audio foundation models, and non-invasive acoustic diagnostics.
Moazzam Shoukat's AI & Speech Research Taxonomy
Mapping 7 peer-reviewed publications across foundation audio architectures, healthcare acoustic sensing, and affective computing.
Transformers in Speech
Overcoming self-attention quadratic complexity O(T^2) in high-sample audio through conformers, subsampling, and chunked inference.
Large Audio Models (LAMs)
Tokenizing continuous soundscapes via neural audio codecs (EnCodec) to power real-time, low-latency generative speech agents.
Internet of Audio Things (IoAuT)
Extracting non-invasive biometric telemetry from respiratory and cardiac acoustics using edge spectrogram neural classifiers.
Affective Computing & Metaverse
Multimodal cross-lingual speech emotion recognition enabling empathetic conversational avatars and synthetic voice defense.
Published Research Papers (7)
Peer-Reviewed & PreprintsTransformers in speech processing: A survey
Comprehensive state-of-the-art review covering transformer architectures across acoustic modeling, automatic speech recognition (ASR), speaker verification, and text-to-speech synthesis.
Sparks of large audio models: A survey and outlook
Investigating emerging foundation audio models, multimodal audio-language representations, zero-shot acoustic comprehension, and the future trajectory of generative audio AI.
Affective computing and the road to an emotionally intelligent metaverse
Explores real-time affective computing interfaces, emotional sentiment extraction from acoustic and physiological cues, and emotionally resonant avatars in virtual worlds.
Transformers in speech processing: Overcoming challenges and paving the future
In-depth analysis of computational efficiency hurdles, low-latency streaming constraints, cross-lingual adaptations, and next-generation transformer paradigms for audio.
Medicine's New Rhythm: Harnessing Acoustic Sensing via the Internet of Audio Things for Healthcare
Groundbreaking evaluation of acoustic sensing, cough/respiratory sound analytics, wearable audio IoT devices, and privacy-preserving clinical diagnostics.
Breaking barriers: Can multilingual foundation models bridge the gap in cross-language speech emotion recognition?
Empirical investigation demonstrating how pretrained multilingual foundation representations bridge the acoustic-emotional divide across low-resource and cross-cultural dialects.
Affective Computing in the Age of GenAI: Balancing Risks and Rewards
Critical inquiry into generative emotional synthesis, ethical boundaries of affective empathy machines, psychological influence, and governance frameworks.
Interested in Research Collaboration?
Open to peer-reviewing, speech AI benchmarks, and acoustic healthcare explorations.