ICON 2026
International Conference on Natural Language Processing (ICON-2026)
Department of Information Technology, Gauhati University, Guwahati, Assam, India
20 -23 December 2026
Keynote Speaker
Department of Information Technology, Gauhati University, Guwahati, Assam, India
20 -23 December 2026
Keynote Speaker
Keynote 1
University of Maryland at College Park
Prof. Dinesh Manocha
Dinesh Manocha is the Paul Chrisman-Iribe Chair in Computer Science & ECE and a Distinguished University Professor at the University of Maryland, College Park. His research interests include virtual environments, physically-based modeling, and robotics. His group has developed numerous software packages that are standard and licensed to over 60 commercial vendors. He has published more than 950 papers & supervised 67 PhD dissertations.
His group has received more than 22 best paper and test-of-time awards at leading conferences in computer graphics, solid modeling, multimedia, VR, and robotics. Manocha is a Fellow of AAAI, AAAS, ACM, IEEE, and NAI, a member of ACM SIGGRAPH and IEEE VR Academies, and a recipient of the Bézier Award from the Solid Modeling Association.
He received the Distinguished Alumni Award from IIT Delhi and the Distinguished Career in Computer Science Award from the Washington Academy of Sciences. He co-founded Impulsonic, a physics-based audio simulation technology developer, which Valve Inc. acquired in November 2016.
Keynote title: General Audio Intelligence (https://sakshi113.github.io/audio_webpage/ )
Abstract:
Audio General Intelligence—the capacity of AI agents to deeply understand and reason about all types of auditory input, including speech, environmental sounds, and music—is crucial for enabling AI to interact seamlessly and naturally with our world. Despite this importance, audio intelligence has traditionally lagged behind advancements in vision and language processing. This gap arises from significant challenges, such as limited datasets, the complexity of audio signals, and a shortage of advanced neural architectures and effective training methodologies explicitly tailored for audio.
In this talk, we give an overview of our work over the last two decades on audio simulation and understanding, and the development of Audio LLMs. This includes earlier works based on audio ray-tracing and scientific solvers. Next, we talk about acceleration and prediction methods based on machine learning. We demonstrate how synthetic simulation methods can be utilized to generate training data and design networks using geometric deep learning methods.
Next, we provide an overview of audio large language models (ALLMs) used for advanced audio perception and complex reasoning. This includes GAMA, which is built with a specialized architecture, optimized audio encoding, and a novel alignment dataset. . Complementing this, we introduced ReCLAP, a state-of-the-art audio-language encoder, and CompA, one of the first projects to tackle compositional reasoning in audio-language models—a critical challenge given the inherently compositional nature of audio. Towards the end, we discuss Audio Flamingo 3 and Music Flamingo, which are our open ALLM models that provide advanced long-audio understanding and reasoning capabilities across speech, sound, and music. We highlight their applications to audio, speech, and music understanding.
Keynote 2
Kavčić-Moura Professor of Language Technologies and Human-Computer Interaction
Language Technologies Institute and Human-Computer Interaction Institute, School of Computer Science, Carnegie Mellon University
Prof. Carolyn Penstein Rosé
Dr. Carolyn Rosé is the Kavčić-Moura Professor of Language Technologies and Human-Computer Interaction in the School of Computer Science at Carnegie Mellon University. Her research advances the concept of Sociotechnical Artificial Intelligence from a highly multidisciplinary perspective, exploring human-AI complementarity and multi-agent human-AI teaming. Her research team pushes the frontier of learnability and generalizability through deeply data-focused explorations of inductive biases, with a current emphasis on abstraction and decomposition. She investigates these issues across multiple problem areas including multimodal conversational process analysis, multimodal document understanding, clinical text processing, knowledge-based question answering, and language models of code. Her research group’s highly interdisciplinary work, published in over 330 peer reviewed publications, is represented in the top venues of 5 fields: namely, Language Technologies, Learning Sciences, Cognitive Science, Educational Technology, and Human-Computer Interaction, with awards in 4 of these fields. She is a Past President and Inaugural Fellow of the International Society of the Learning Sciences, Senior member of IEEE, Founding Chair of the International Alliance to Advance Learning in the Digital Era. She also serves as a 2020-2021 AAAS Leshner Leadership Institute Fellow for Public Engagement with Science (AI Cohort).
Keynote title: Optimizing Towards a Powerful Synergy within the Design Space of Collaborative LLM Agents
Abstract:
This talk reports on continuing work developing AI agents serving as first class collaborators rather than either lone actors or subservient helpers, appliances, or tools. Though the area of human-AI collaboration is relatively young, benchmarks for LLM contribution to this space are beginning to emerge. In contrast to work on LLM benchmarks that implicitly cast AI agents as lone actors, and therefore implicitly casting humans as abdicating to agents rather than maintaining their meaningful agentic role, this work seeks to optimize role definitions both for AI agents and for humans within the space of multi-agent human-AI collaboration. This work breaks new ground in terms of task type (e.g., messy, non-routine tasks without a well-established break-down or formally specified action and state space) and roles (e.g., tightly coupled interactive work with fluid role taking and shared leadership rather than divide-and-conquer approaches and fixed roles). An iterative 3-stage optimization process involving a multifaceted post-training pipeline, simulation studies, and user studies is used to efficiently navigate the enormous design space. The goal is to afford collaboration between humans and AI agents that leverages the uniquely human ability to approach problems through abstraction and decomposition alongside model capabilities for pattern-based retrieval from massive storehouses of examples that dwarf the experience any human could bring to the process.
Relevant Links
Kotalwar, N., Das, R., & Rose, C. (2026). Searching for Synergy in Shared Workspace Human-AI Collaboration. arXiv preprint arXiv:2606.18413.
Professor Pushpak Bhattacharyya Memorial Keynote Talk
Professor-Emeritus IIIT Hyderabad, India
Prof. Rajeev Sangal
Prof Rajeev Sangal along with late Dr Vineet Chaitanya has developed Computational Paninian Grammar (CPG) framework which is linguistically elegant and computationally efficient. It forms the basis of most work on Indian language processing today. He has written 5 books and numerous research papers.
He set up Language Technologies Research Centre at IIIT Hyderabad as its founder-head in 2002. It is the largest academic centre for computational linguistics (CL) and Natural Language Processing research in this region of the world.
He was a member of the editorial board of Computational Linguistics in 2004-06. He has been a member of editorial boards of journals: Machine Translation, Natural Language Engineering, Trans. of Asian Language Processing, CSI Transactions, etc. He reviews for major conferences such as ACL, COLING, IJCNLP, ICON, etc. He was the founding President of NLP Association India from 2002 to 2021, and helped
establish a quality annual conference ICON.
He was the architect of Bhashini, the first AI mission of India, and served as the founding Chairman of its Executive Committee.
Keynote title: Discourse in Machine Translation
Abstract:
Machine translation (MT) systems have to learn to handle 'discourse', and must not be limited to handling single sentence at a time. Handling of discourse segments in machine translation helps in producing translation output with much better flow and coherence.
Department of Information Technology, Gauhati University, Guwahati, Assam, India