Abstract
Most AutoML tools are black-box tools. They offer no code/low code tools (UI/simple APIs) for practitioners to get started quickly. While this helps beginners, most experienced data scientists/ML practitioners often need more control. Building a predictive model is an iterative process, so such restricted behavior of AutoML tools leads to the limited use of it. Keeping the real-life data scientist in mind, we created a programming model called “Gradual AutoML”. It borrows some concepts from functional programming and addresses the entire spectrum of controlled automation. Gradual AutoML allows the data scientist to be in the driver’s seat and use AutoML for assisted driving.
This talk will cover the basics of AutoML and then present Lale (https://github.com/IBM/lale), an open-source scikit-learn compatible AutoML library which implements Gradual AutoML. It will include usage examples and code showing how ML practitioners can control certain choices and employ AutoML to do the rest. I will also briefly share how to use Lale for AutoML with imbalance correction, computation of fairness metrics and bias mitigation. The talk assumes some familiarity with the Python ML ecosystem, but many of the concepts apply to the general AutoML framework
Interview
I work on AI research. The session I'm going to conduct is on AutoML or AutoAI. Right now, I'm working on AutoAI with foundation models in mind, which are the large language models, the latest in AI.
Most of the commercial or open-source AutoML tools today are black-box tools. For data scientists or ML practitioners who want to use the optimization techniques that AutoML provides, they have a very black-box interface. They can give their data, tasks, and maybe some other hyperparameters, but that's it. What we want to achieve is to give more control to the data scientists, so they can inject their domain knowledge and intuition into the AutoAI process. Instead of being a black-box tool, they can have control and provide algorithm choices, hyperparameters, or even the search space. This way, they can try out an iterative process for AutoAI.
Ideally, I think I would expect them to have some knowledge of using ML, and it would be even better if they have knowledge of the Python ecosystem for machine learning, which includes open-source libraries like Pandas and scikit-learn. If they have used AutoAI, that's great, but if they haven't, I would cover the basics of what it means to take assistance from AutoAI.
Yes, I would like them to walk away with the understanding that AutoAI is not daunting or a black-box. They can control a lot of things and even perform complex tasks with it. They can drive how it searches and uses optimization. If they want to tackle complex use cases like imbalance correction or fairness mitigation, that's all possible. They should use the right tool to leverage that.
Topics
QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
ML in Practice Hosted by Sid Anand Chief Architect and Head of Engineering @DatazoomFrom the same track
Thursday 15 June
10:35 Dumbo / Navy Yard Session AI/ML PostgresML: Leveraging Postgres as a Vector Database for AI Montana Low Machine Learning w/ PostgresML With the growing importance of AI and machine learning in modern applications, data scientists and developers are constantly exploring new and efficient ways to store and analyze large amounts of data. 11:50 Dumbo / Navy Yard Session Search Needle in a 930M Member Haystack: People Search AI @LinkedIn Mathew Teoh Machine Learning @ LinkedIn LinkedIn's search functionality is one of its oldest capabilities, allowing members to search for people they know, or to discover new connections. 13:40 Dumbo / Navy Yard Session AI/ML Going Beyond the Case of Black Box AutoML Kiran Kate Senior Technical Staff Member @IBM Research Most AutoML tools are black-box tools. They offer no code/low code tools (UI/simple APIs) for practitioners to get started quickly. While this helps beginners, most experienced data scientists/ML practitioners often need more control. 14:55 Dumbo / Navy Yard Session ML in Practice Back to Basics: Scalable, Portable ML in Pure SQL Evan Miller Principal Statistics Engineer @Eppo (Creator of Evan's Awesome A/B Tools) Redshift has SageMaker. BigQuery begat BigML. Spark birthed Databricks. Every data warehouse is tightly coupled to a particular ML stack. 16:10 Dumbo / Navy Yard Session LLMs in the Real World: Structuring Text with Declarative NLP Adam Azzam AI Product Lead @Prefect Building machine learning pipelines to extract structured data from unstructured text is a popular problem within an unpopular development lifecycle.