Minimal Data Cleaning for Model Training by MinPrep
PVLDB 2026 (to appear), 2026
MinPrep supports minimal data cleaning for supervised model training, helping users focus cleaning effort on the data repairs that matter for downstream model quality.
PVLDB 2026 (to appear), 2026
MinPrep supports minimal data cleaning for supervised model training, helping users focus cleaning effort on the data repairs that matter for downstream model quality.
arXiv cs.LG/cs.AI, 2026
This paper models interactions between users and ML systems whose incentives may not be aligned. It proposes a game-theoretic framework and scalable algorithms that help users benefit from ML systems while reducing biased or manipulative actions.
arXiv cs.DB, 2026
This paper studies how users can query data sources that may intentionally return biased answers because their incentives do not align with the user’s information needs. It proposes algorithms to detect biased information and reformulate queries to recover more relevant results.
NeurIPS Reliable ML from Unreliable Data Workshop, 2025
This work studies how to repair incomplete training data only where it matters for downstream learning. Instead of imputing every missing value, it focuses data preparation effort on the minimal repairs needed to learn accurate models.
arXiv cs.LG, 2025
Missing data often exists in real-world datasets, requiring significant time and effort for imputation to learn accurate machine learning (ML) models. In this paper, we demonstrate that imputing all missing values is not always necessary to achieve an accurate ML model.
IUI, 2025
This paper explores potential parallels between the evolution of users’ interactions with visualization tools during data exploration and assumptions made in popular online learning techniques. Through a series of empirical analyses, we seek to answer the question: What are the best learning methods for modeling shifts in users’ data focus during EVA?
ICDE Lightning Talk, 2024
This paper studies how users learn while exploring data interactively and how data systems can better account for evolving user understanding during exploratory visual analysis.
SIGMOD, 2024
Visualization Recommendation Systems help users discover important insights during data exploration. These systems should understand users’ exploration.
SIGMOD, 2024
Real-world data is often incomplete and contains missing values. To train accurate models over real-world datasets, users need to spend a substantial amount of time and resources.
AAAI, 2024
Exploratory visual analysis (EVA) is an essential stage of the data science pipeline, where users often lack clear analysis goals.
Conference, 2022
A known problem with training deep neural networks, mostly parameterized by connection weights at each layer, is that of finding an appropriate model complexity under the empirical risk minimization setting.
, 2022
This project proposes a deep learning-based approach that simulates the user-feedback loop in entity matching.
ICAART, 2021
The agriculture and farming industry plays a vital role in the economy. However, the importance of agriculture cannot be fully quantified in terms of its economic profit.