LUIAS Summer Workshop 2025

Workshop Banner

Objective

The Summer Workshop facilitates collaboration among worldwide data science researchers to exchange research and professional insights on AI science and engineering and to foster diverse partnership opportunities within the industry. This year, our theme focuses on AI Engineering across various sectors, including transportation, energy, healthcare, engineering design, and manufacturing.

Speakers

Abstracts

In a smart manufacturing system, a large number of sensors are installed to monitor machine status, process variables, product quality, and the overall system performance.  It is always a challenging problem on how to analyze those massive amounts of data effectively for cost reduction and quality improvements in all manufacturing companies. This presentation will discuss research opportunities, challenges, and advancements in this important research area, especially how machine learning concepts and algorithms can be used to solve challenging quality improvement problems. Examples of ongoing research projects will be used to articulate the frontiers of this research area. All examples come from real data and problem in industrial production systems. This presentation will emphasize the motivations of these research undertakings: challenges to be overcome, new methods that were developed, validation/implementation undertook, as well as the potential impacts. 

Spatial transcriptomics (ST) has advanced our understanding of tissue regionalization by enabling the visualization of gene expression within whole-tissue sections, but current approaches remain plagued by the challenge of achieving single-cell resolution without sacrificing whole-genome coverage. In this talk, Prof. Kaibo Liu will discuss a recent publication from his research group in Nature Methods, conducted in close collaboration with St. Jude Children’s Research Hospital. He will present Spotiphy (spot imager with pseudo single-cell-resolution histology), a computational toolkit that transforms sequencing-based ST data into single-cell-resolved whole-transcriptome images. Spotiphy delivers the most precise cellular proportions in extensive benchmarking evaluations. Spotiphy-derived inferred single-cell profiles reveal astrocyte and disease-associated microglia regional specifications in Alzheimer’s disease and healthy mouse brains. Spotiphy identifies multiple spatial domains and alterations in tumor–tumor microenvironment interactions in human breast ST data. Spotiphy bridges the information gap and enables visualization of cell localization and transcriptomic profiles throughout entire sections, offering highly informative outputs and an innovative spatial analysis pipeline for exploring complex biological systems. 

The introduction of AlphaFold has revolutionized the task of protein structure prediction from a given sequence of amino acids; the groundbreaking contribution of AlphaFold was recognized by the 2024 Nobel Prize in Chemistry. As a deep-learning based method, AlphaFold was trained from the publicly available Protein Data Bank (PDB), a database of known protein structures. An inherent limitation of AlphaFold is that its prediction can only give a static structure, whereas in reality, the structures of proteins are dynamic and can change in response to their environment or binding partners, with significant biological consequences. In this talk, we focus on enhancing and diversifying protein structure prediction using AlphaFold. Through a principled statistical sampling framework, we significantly expand AlphaFold’s capabilities, enabling it to explore a broader conformational space. Key methodologies involve modifying the multiple sequence alignment (MSA) and template inputs to encourage AlphaFold to explore different conformations, thereby increasing structural diversity. This is achieved in particular through a sequential sampling approach, which allows for the creation of multiple diverse MSAs, broadening the conformational possibilities that AlphaFold can investigate. We will illustrate the capabilities of the sequential sampling approach through examples.

We consider the problem of learning a continuous probability density function from data, a fundamental problem in statistics known as density estimation. It also arises in distributionally robust optimization (DRO), where the goal is to find the worst-case distribution to represent scenario departure from observations. Such a problem is known to be hard in high dimensions and incurs a significant computational challenge. In this talk, I will present a machine learning approach to tackle these challenges, leveraging recent advances in generative models, which have become popular recently due to their competitive performance in high-dimensional data. We show that flow-based generative models provide a flexible computational framework for such problems, and they can be cast as particle-based iterative algorithms in probability space with the Wasserstein metric. Based on this simple and general framework, we can prove the convergence of the iterative algorithm and show the generative guarantee, meaning identify suitable conditions, the learned density is close to the true distribution. We demonstrate the utility of this framework for density estimation, distributionally robust optimization (DRO), and applications to adversarial learning.

Quantitative models are used extensively in banks and financial institutions to assess and manage risks posed by their products and services. Common data-centric examples include credit decisions, loss forecasting, fraud detection, and complaints classification, the last using natural language processing. Machine learning (ML) algorithms have become increasingly common due to their predictive performance, but the black-box nature of the algorithms have restricted widespread use in large institutions that are regulated. I will provide an overview of the applications, the need for explainability, and recent research done in this direction. The use of Generative AI (GenAI) and large language models (LLMs) is being explored for efficiency gains, such as document automation, compliance/policy checks, etc. I will give an overview of their uses and challenges in these contexts.

Owing to the high level of complexity of AI models, it is often difficult, if not impossible, to understand how they work. This has given rise to the research field of explainable AI (XAI), interpretable AI, and explainable machine learning (XML), where the goal is to make the AI and ML algorithms transparent, interpretable, and explainable. Proposed solutions include (1) partial dependency plots and Shapley values that show the marginal effects of each input feature, (2) feature importance measures that estimate the importance of a feature, and (3) approximating an AI model with locally interpretable models. The goal is even harder to achieve if there are missing values in the data used to train the AI models, because explanations should include how the missing values are dealt with. This talk shows how the GUIDE regression tree algorithm can solve these problems.

There has been significant developments of large and complex engineered infrastructure systems such as telecommunication networks, power grids, transportation infrastructure, healthcare delivery systems, information technology, financial systems and supply chain systems. Failures of such systems may result in cascading damages as well as significant interruptions of their services. In the last two decades, these systems have experienced increasing natural and manmade hazards, which have caused significant disruptions of their functions. Moreover, these failures as well as the inherent failures of the systems and their repairs affect the overall system’s availability.  This talk presents approaches for resilience quantification and system’s availability under these conditions.  

Maintenance optimization for complex systems is an increasing critical issue in manufacturing industries including automobiles and semiconductors. Using IoT and smart censors, engineers aim to decide proper maintenance time points or intervals via health indicators representing system conditions. In this seminar, I introduce prognostic health management (PHM) and predictive maintenance (PdM) via off-line and on-line data for complex systems. Using off-line data, I present statistical models (e.g., nonhomogeneous Poisson process (NHPP), frailty models) for repairable systems. For PHM, I introduce a general five-stage process for PHM and PdM. I also present condition based maintenance policy using signal processing and statistical process control techniques, based on on-line sensor data. Finally, I present several real case studies for PHM and PdM in power plants and automobiles.

There are two main approaches to supply chain modeling: the bottom-up and top-down approaches. The top-down approach is the more traditional approach to supply chain optimization, which focuses on the design, planning, and scheduling of supply chains. On the other hand, the bottomup approach targets the transactional processes that are present in Enterprise Resource Planning (ERP) systems. In this presentation we address the development of a holistic approach to supply chain management, requiring a paradigm shift from operation-based decision support systems to integrated decision frameworks that account for the different areas (e.g., accounting, research and development, sustainability) and flows (material, financial, and information) associated with supply chains. As a first step, a hybrid optimization/machine learning framework is presented to integrate the material flows in a chemical batch plant with the information flows in the order fulfillment or order-to-cash (OTC) supply chain process. We next describe a more comprehensive approach, which incorporates manufacturing scheduling models in the OTC process model providing a more complete and accurate view of the supply chain by accounting for both material and information flows. A stochastic discrete-event simulation framework is used to dynamically model the system behavior. Optimization events are triggered each time a new order enters the system, at which time a comprehensive mixed-integer linear State-Task Network optimization model is called to schedule both the order processing steps and the plant operations. Whenever an optimization event is completed, updated order priorities and queue assignments are passed to the transactional queues in the discrete event simulation along with the updated production schedule, which may be modified with reinforcement learning techniques. It is shown that purely transactional models can yield suboptimal results because these models do not account for synergies in the manufacturing plant arising from co-production, which allows reducing the order fulfillment lead times. On the other hand, purely physical models can result in schedules that are infeasible. This occurs because the models do not account for bottlenecks in the transactional process, which affect raw material availability at the plant. Thus, the integrated approach finds a solution that is more profitable than the purely transactional schedule, and corrects for the infeasibilities in the purely physical schedule.

Historically, development of rigorous maintenance planning models has generally involved the application of reliability and maintenance probabilistic models and theory combined with formal analytical tools associated with operations research and mathematical programming. These models, while effective, tend to produce static and non-changing strategies that so not reflect changing conditions. Furthermore, they are based on population characteristics and do not adequately reflect individual differences of units within the population or system. Alternatively, predictive maintenance models which embraced machine learning can dynamically predict a remaining useful life or RUL. These models do specifically reflect individual units within the system and can also adapt to changing conditions. However, their true effectiveness also requires a meaningful decision-rule on when to take action and what action, i.e., either replace or repair or dynamically reduce work-load to compensate for anticipated degradation. To be truly effective, a combination of these philosophies is needed. The usefulness of predictive maintenance decision rules require the three-way integration or intersection of machine learning, reliability/maintenance theory and operations research. In this talk, we will summarize different approaches to preventive or predictive maintenance models, discuss their relative advantages and disadvantages, highlight a few notable examples to demonstrate this three-way integration, and finally present future research challenges.

The development of cyber-physical systems with multiple sensor/actuator components and feedback loops has given rise to advanced automation applications, including energy and power, intelligent transportation, water systems, manufacturing, etc. Traditionally, feedback control has focused on enhancing the tracking and robustness performance of the closed-loop system; however, as cyber-physical systems become more complex and interconnected and more interdependent, there is a need to refocus our attention not only on performance but also on the resilience of cyber-physical systems. In situations of unexpected events and faults, artificial intelligence and machine learning can play a key role in improving the fault tolerance of cyber-physical systems and preventing serious degradation or a catastrophic system failure. The goal of this presentation is to provide insight into the design and analysis of intelligent monitoring methods for cyber-physical systems, which will ultimately lead to more resilient societies.

Process monitoring plays a crucial role in various manufacturing and service industries. Control charts have been widely used for this purpose because they provide a visual representation of process performance, making interpretation straightforward. As a result, engineers without a statistical background can easily understand them. However, control charts have limitations because they rely on certain statistical assumptions, making them less effective in handling complex situations commonly found in modern manufacturing processes. Recently, machine learning and deep learning-based techniques have gained popularity in process monitoring, often under the term "anomaly detection." In this talk, I will discuss the evolution of process monitoring, from traditional control charts to the latest deep learning-based approaches.

Artificial Intelligence (AI) is clearly one of the hottest subjects these days. Basically, AI employs a huge number of inputs (training data), super-efficient computer power/memory, and smart algorithms to perform its intelligence. In contrast, Biological Intelligence (BI) is a natural intelligence that requires very little or even no input. This talk will first discuss the fundamental issue of input (training data) for AI.  After all, not-so-informative inputs (even if they are huge) will result in a not-so-intelligent AI. Specifically, three issues will be discussed: (1) input bias, (2) data right vs. right data, and (3) sample vs. population. Finally, the importance of Statistical Intelligence (SI) will be introduced. SI is somehow in between AI and BI. It employs important sample data, solid theoretically proven statistical inference/models, and natural intelligence. In my view, AI will become more and more powerful in many senses, but it will never replace BI.  After all, it is said that “The truth is stranger than fiction, because fiction must make sense.”   The ultimate goal of this study is to find out “how can humans use AI, BI, and SI together to do things better.”

The widespread use of AI will have far-reaching effects on the scientific publication process. Ways that AI can benefit authors will be discussed in this presentation along with the effect of AI on the review process. Some of the approaches used to attack the scientific publication process will be covered and illustrated with examples. In this respect the use of AI has the potential to lead to an increasing number of submissions of fake papers produced and sold by paper mills. My experience with authors’ use of AI as editor of the Taylor & Francis journal Quality Engineering will be shared with ample time left for open discussion where participants can share their thoughts and experiences.

Co-Chairs