Bite-sized

Day 1: Thursday, 12 September

-

Future Data Services 1

Session convener: Vassilis Routsis, Xan Morice Atkinson

Session 1.A: Developing systems and methodologies to enhance discovery and retrieval of complex data using LLMs
Large Language Models and Generative AI in general are attracting significant attention and finding applications in various environments. This session will explore how this technology can enhance data accessibility for research purposes. We will discuss the aims, objectives, and methodologies of CORDIAL-AI, a project designed to help researchers discover and retrieve information from complex structured datasets, such as census origin and destination data, using the power of Generative AI. We will present some of the latest advancements in this field, from fine-tuning models to creating efficient pipelines for user interaction, and share our approach and methodologies. The project utilises open-source models and software to provide an in-house solution. We will discuss the potentials, challenges, benefits, and limitations of this technology and our approach. Attendees will be invited to share their ideas, concerns, experiences, or any input related to the application of this technology for data discovery and retrieval. While the session will touch on several technical aspects, advanced knowledge in the field is not required.

Session 1.B: Data discovery made easy: Applying ML to a diverse social science database
Why can I "talk" to a database now and how exactly did we as a society get here? AI, and specifically Large Language Models (LLMs), are a major disruptive technological advancement that have, and will, drastically change society for both better and worse.In this session, we cover the evolution of the academic landscape in AI methods - including the new research methods that we utilise in our project called “Data Discovery Made Easy” (DDME). We explain how we are currently using open-source software built around LLMs to bridge the gap between our geospatial database and users. We also discuss the potential usefulness of advanced, locally-hosted, open-weight models, and potential scaling methods for tool production and deployment.This new type of LLM-based data discovery method raises many questions: How and why does it work? Where does it fail? How is it tested? What are “hallucinations” in the context of LLMs? Why do they happen? Why are they important to us and the DDME project? Can they be stopped? Where will the AI research field go next? What about ethics? This is a session to learn about some of these topics together.