Research

Data Lake Research and Development Center

Realizing Advanced Data Management for the Generative AI Era

Since the advent of ChatGPT, generative AI has had a profound impact on how society functions, and it is expected not only to boost productivity but also to accelerate scientific research and innovation. Applying generative AI to advanced, specialized purposes requires feeding diverse domain data into a base AI model to build a more refined model. Data Lake Research and Development Center (DLRD) works to build a data infrastructure that comprehensively manages diverse training data and AI models — together with their provenance and conditions of use — for an era of generative AI in which multiple institutions make mutual use of one another’s proprietary data and AI models.

Provide Integrated AI Models and Training Data via a Secure Environment

A data lake is a data infrastructure architecture that stores structured data (such as numeric values and strings) and unstructured data (such as documents and images) in a unified way, enabling flexible, exploratory data analysis. DLRD’s data infrastructure extends management to AI models that have been trained on the data, providing a research and development platform for the continuous advancement of AI models — through the addition and updating of training data and the refinement of learning methods — and for analyzing AI model behavior by tracing back through training data. Our main efforts are:

  • Continuously collecting diverse training data while managing its provenance and usage rights while also building vector indexes on them to enable retroactive search.
  • Establishing an advanced AI model operational practice that manages provenance information — including training data and the training process — for each version of an AI model.
  • Building a robust system with advanced security, providing a safe data access environment for a wide range of cutting-edge research fields as an inter-university research institution.
  • Identifying legal issues (such as those under the Act on the Protection of Personal Information and the Copyright Act) and socio-ethical issues (such as bioethics) in the collection and use of data and incorporating these into the management of training data and AI models.

Building and Operating a Medical Data Lake to Promote Medical LLM/LMM R&D

In Japan’s medical field, both academic considerations (promoting research) and social considerations (ensuring the appropriate provision of medical technology) call for the sharing of medical data and the use of generative AI. DLRD is realizing a data infrastructure for the medical field, participating in the construction and operation of a medical data lake as part of SIP (Cross-ministerial Strategic Innovation Promotion Program) Phase 3’s “Building an Integrated Healthcare System.” We have implemented a large-scale medical data lake with 10 petabytes of storage and a governance mechanism for large-scale medical LLM/LMM models, and are systematically collecting training data consisting of clinical medical data, medical knowledge bases, and large-scale Japanese web data. We are also managing the appropriate use of medical data and medical LLMs/LMMs using a medical data governance database.

Members

Takayuki Tamura
Director, Project Professor
Miyuki Nakano
Executive Director, ROIS
Kento Aida
Professor
Hiroki Takakura
Professor
Masakazu Hayashi
Project Researcher
Yusuke Matsui
Project Researcher