As part of the 5th Midwest Healthcare Conference, we are excited to host a data challenge that invites participants to push the frontiers of large language model (LLM) development in service of a critical mission: improving how AI understands and responds to real-world medical questions. This competition fosters innovation at the intersection of artificial intelligence, clinical reasoning, and digital health, offering a unique opportunity to contribute to transformative research with real-world impact.
A large language model (LLM) is a type of deep learning algorithm designed to understand, summarize, translate, predict, and generate text by learning from vast amounts of data.
LLMs represent one of the most impactful uses of transformer models. Their capabilities go far beyond processing human languages; beyond enhancing natural language processing tools such as translators, chatbots, and virtual assistants, LLMs are increasingly used in healthcare, coding, and human-computer interaction. As these models become more integrated into decision-making and service delivery, their influence on research, communication, and everyday human-computer interaction continues to grow rapidly.
In healthcare, LLMs are transforming clinical decision support, streamlining documentation, and improving patient engagement through intelligent interfaces. Their ability to analyze unstructured medical text and integrate diverse sources of health data gives them the potential to significantly enhance diagnostic accuracy, alleviate administrative burdens, and support more personalized care, among other impactful applications.
The competition invites participants to develop large language models (LLMs) specifically trained to respond to patient-initiated clinical questions with clarity, accuracy, and contextual sensitivity. Unlike general-purpose LLMs trained primarily on web-based text, this contest emphasizes the use of diverse and domain-relevant data sources, including medical literature, clinical notes, structured datasets, and curated question-answer pairs, to build models with specialized healthcare competencies.
Participants are encouraged to explore innovative strategies in data selection, model architecture, and fine-tuning. The aim is not only to improve factual correctness and safety in medical responses, but also to design models that can handle nuanced patient language, simulate clinical reasoning, express uncertainty appropriately, and communicate with empathy.
The challenge is to develop cutting-edge LLMs specifically designed to answer patient-initiated clinical questions with accuracy, clarity, and empathy. The ultimate goal is to enhance LLM capabilities in providing safe, context-aware, and clinically relevant responses to patients seeking medical guidance.
Participants are not expected to train a language model from scratch. Instead, they are encouraged to build upon existing pretrained models and adapt them to healthcare using innovative and high-quality data sources. Publicly available datasets and domain-relevant corpora may be used subject to applicable licensing, privacy, and ethical requirements.
Models developed for the competition are expected to be made publicly available under an appropriate open-source license upon completion. Participants are strongly encouraged to implement safeguards against misinformation, including hallucination detection, uncertainty estimation, and expert adjudication.
Models are assessed on the innovation of their design, training methodology, and data strategy.
Responses are assessed for accuracy, relevance, clarity, depth of contextual understanding, and appropriate communication.
The competition is hosted on the Kaggle platform. Participants may register as teams of 1–4 members and use the platform to build, test, and refine their models.
The top three teams will be invited to join the organizing team as co-authors on a planned peer-reviewed publication and to present at the conference. Awards include First Prize $1500, Second Prize $500, and Third Prize $250.
Main Contact: midwest.data.competition@gmail.com
Organizers include Mehmet Eren Ahsen and collaborators from the University of Illinois.