Product: AI Safety Programs Role: Technical Program Manager Year: 2022-2025
Description: The AWS AI Safety team was responsible for the technical and scientific support of classic and generative AI development across AWS. In this role I partnered with applied scientists, ML engineers and product owners to manage end-to-end AI evaluation and safety programs.
WHAT'S THE PROBLEM BEING SOLVED?
A) AI services can generate biased or inaccurate responses. Businesses can't trust and confidently incorporate those services into their operations without transparent, scientifically verifiable documentation that identifies these risks.
B) AI Services are only as fair and unbiased as their training data. Independent safety evaluations from high quality expert datasets are needed to help them improve.
WHO ARE WE SOLVING IT FOR?
A) Business customers deciding whether to incorporate AWS AI services into their critical applications.
B) AI Service developers who need scientific and technical support to improve their services based on the results of the safety evaluations.
HOW WE SOLVED IT
The solution entails the four step process below which comprises a technical program. Risk assessments, and evaluation theses on AI models and services result in confidential information and documentation that will not be addressed in this case study.
I used Gantt charts like the example below to track the production of multiple datasets, and the science and engineering resources to support production.
Data Vendor Selection & Management
High quality evaluation datasets required contracting and managing data vendors from all over the world. I managed the drafting and execution of vendor contracts from six to seven figures.
Production Workflow Management
Dataset production is a time consuming process requiring the managed coordination of multiple programs across multiple roles. Bottlenecks ultimately delay model production and release schedules. There was constant optimization and iteration of the workflows.
Annotation & QA Management
Where necessary I managed the human labeling of raw data, which is a process that A) entails the design and development of annotation tools, B) the testing, selection and training of human annotators, and C) the verification and quality assurance validation of the labels. The screenshots below are from my dashboard of annotator qualifications, task tracking and agreement rate for a speech collection dataset.
Evaluations
Completed datasets are used in evaluations conducted by applied scientists and are very comprehensive and confidential. The example below shows the relationship between a completed speech dataset production and the resulting evaluation on one set of pair-wise attributes.
WHAT WE LAUNCHED
AWS AI Service Cards
In the fall of 2022 we launched AWS AI Service Cards to the public via the aws.amazon.com website. Cards for generative AI models and services were introduced at AWS re:Invent in 2023 by the Sr. VP of AI. I managed the programs behind all of the service cards for speech, language and health AI, while providing technical support for our in-house LLM service cards. Some of the most challenging cards I managed included:
Amazon Q
Business, AWS
HealthScribe and Amazon
Transcribe Toxicity Detection.
In addition to the AI service cards my team produced and published guidance on our evaluation process and dataset production best practices. The AWS Responsible AI web page experienced high traffic as the customer demand rose for AI safety documentation after 2023.
Impact
Normally this is where I would share some metrics that validated the success of the solution and launch. While I can't share any of that data, I can share two things that support the impact. The first is that this program directly contributed to the framework of AI system management that helped AWS become the first hyper-scale cloud provider to receive ISO 42001 certification for multiple AI services. The second is a social media post demonstrating that my team's work had an impact on earning customer trust, an invaluable metric.
–