How robots can adapt to new tasks — quickly

New approach to meta-reinforcement learning minimizes the need for costly interactions with the environment.

Reinforcement learning (RL) is a technique in which an AI agent interacts with an environment and learns a policy based on the rewards that it receives during this interaction. Progress in RL has been dramatically demonstrated by human-level performance on games like Atari. The key to this progress was generating large amounts of data using game simulators.

There are two hurdles in translating this progress into real-world applications such as assembly-line robots or robots that help the elderly in their homes. First, robots are complex and fragile; learning by taking random actions could damage the robot or its surroundings.

Second, the environment in which a robot operates is often different from the one it was trained for. A self-driving car, for instance, might have to work in a different part of the city from the one in which it was trained. How can we build learning machines that can handle new scenarios?

In a paper that we will present at the International Conference on Learning Representations, we describe a new reinforcement learning algorithm named MQL (for meta-Q-learning) that enables an AI agent to quickly adapt to new variations of familiar tasks.

Learning to learn

With MQL, as with other “meta-learning” algorithms, an agent is trained on a large number of related tasks — e.g., how to pick up objects of different shapes — and then tested on how well it learns new variations of those tasks.

MQL has two key differences. The first is that during training, the agent learns to compute a context variable specific to each task. This enables it to learn different models for different tasks: picking up a coffee cup, for instance, is much different from picking up a soccer ball.

Second, during testing, MQL uses a statistical technique called propensity estimation to search its training data for past interactions that look similar to those from the new task it’s learning. This allows MQL to adapt to the new task with minimal interactions.

Credit: Stacy Reilly

Consider the robot above, which wants to learn to pick up objects. In the RL framework, the robot would try to pick up the objects; it would get a reward every time it successfully picked one up and a penalty if it dropped it.

Over repeated trials, the robot learns a policy that enables it to pick up all the objects in the training set. It is likely to do better, however, if that policy includes different interaction models for different objects.

This is the first key idea behind MQL: the robot learns a context that differentiates the model for the mug from that of the soccer ball. MQL uses a gated-recurrent-unit (GRU) neural network to create a representation of the task, and the system as a whole is conditioned on that representation.

Reusing data

The context helps the system predict a model for handling a new task — say, picking up a bottle of water. Adapting that model, however, can still require a large number of training samples. This brings in the second key component of MQL: its use of propensity estimation.

A propensity score indicates the odds that a given sample came from either of two distributions. MQL uses propensity estimation to decide which parts of the training data are close to the test task data: picking up a bottle, for instance, is closer to picking up a mug than to picking up a soccer ball. The model can then sample from the relevant training data, augmenting the data from the new task in order to adapt more efficiently.

We also used propensity estimation in our “P3O: Policy-on Policy-off Policy Optimization”, which we presented at the Conference on Uncertainty in Artificial Intelligence (UAI) in July 2019. There, too, the technique helped reduce the number of samples required to train reinforcement learning algorithms.

As AI systems tackle larger and larger sets of applications, the amount of data available for training begins to feel small. Techniques like MQL are a way to bootstrap the learning of new tasks from existing data and dramatically reduce the data requirements for training AI systems.

Research areas

Related content

US, WA, Bellevue
We are designing the future. If you are in quest of an iterative fast-paced environment, where you can drive innovation through scientific inquiry, and provide tangible benefit to hundreds of thousands of our associates worldwide, this is your opportunity. Come work on the Amazon Worldwide Fulfillment Design & Engineering Team! We are looking for an exceptional Principal Reliability Engineer, someone that is excited to work on complex real-world challenges for which a comprehensive scientific approach is necessary to drive solutions. Your investigations will include identifying and mitigating reliability risks resulting in design and implementation of safe and reliable fulfillment center designs. We are seeking a Principal Reliability Scientist who will be responsible to lead reserach specific to reliability for Amazon Fulfillment Facilities. Key job responsibilities a. Employing a Failure based approach to develop and implement both analytical and empirical methods for identifying and evaluating reliability risks at the stages of facility design, MHE selection, and deployment. b. Lead research activities related Design for Reliability (DfR) initiatives for a warehouse, Develop predictive modelling techniques, and test plans. c. Lead AI-powered reliability solutions that combine data, machine learning, and domain knowledge about asset operations to help anticipate reliability risks, increase uptime, manage CapEx and OpEx costs, and improve productivity. d. Identifying gaps in current Reliability standards and guidelines, and lead comprehensive research to redefine “industry best practices” based on solid scientific foundations. e. Continuously strive to gain in-depth knowledge of your profession, as well as branch out to learn about intersecting fields, such as robotics and mechatronics. f. Travelling to our various sites to perform thorough assessments and gain in-depth operational feedback, approximately 25%-50% of the time. g. Regular communication with Executive management regarding status, risks and operational program metrics.
US, WA, Seattle
Interested in helping build Prime's content and offer personalization system to drive huge business impact on millions of customers? Join our team of Scientists and Engineers developing algorithms to adaptively generate, optimize, and personalize the customer experience with Amazon Prime. This includes identifying who our customers are and providing them with personalized relevant content. As an ML scientist, you will partner directly with product owners to intake, build, and directly apply your modeling solutions. There are numerous scientific and technical challenges you will get to tackle in this role, such as deep learning and reinforcement learning, and their application to various types of contextual, multi-step optimization of the customer journey. We employ techniques from supervised learning, multi-armed bandits, optimization, and RL - while this role is focused on the space of discriminative and generative recommender systems. As the central science team within Prime, our expertise gets routinely called upon to weigh in on a variety of topics. We also emphasize the need and value of scientific research and have developed a strong publication and patent record (internally/externally) which you will be a part of. You will also utilize and be exposed to the latest in ML technologies and infrastructure: AWS technologies (EMR/Spark, Redshift, Sagemaker, DynamoDB, S3, ...), various ML algorithms and techniques (Random Forests, Neural Networks, supervised/unsupervised/semi-supervised/reinforcement learning, LLMs), and statistical modeling techniques. Major responsibilities - Build and develop machine learning models and supporting infrastructure at TB scale, in coordination with software engineering teams. - Leverage Bandits, Supervised Learning, and Reinforcement Learning for Contextual Recommendation and Optimization Systems. - Develop offline policy estimation tools and integrate with reporting systems. - Establish scalable, efficient, automated processes for large scale data analyses, model development, model validation and model implementation. - Analyze and extract relevant information from large amounts of Amazon’s historical business data to help automate and optimize key processes. - Work closely with the business to understand their problem space, identify the opportunities and formulate the problems. - Use machine learning, data mining, statistical techniques and others to create actionable, meaningful, and scalable solutions for the business problems. - Design, develop and evaluate highly innovative models and statistical approaches to understand and predict customer behavior and to solve business problems.
US, CA, Santa Monica
We are the Sponsored Products - Marketplace Intelligence (MI) team. We are looking for a Data Scientist to help build production ML and bandit solutions to customize the search experience. We determine which ads to show in Amazon search, where to place them, how many ads to place, and to which customers. This helps shoppers discover new products while helping advertisers put their products in front of the right customers, aligning shoppers’, advertisers’, and Amazon’s interests. To do this, we apply a broad range of machine learning, causal inference, and optimization techniques to continuously explore, learn, and optimize the allocation and ranking of ads on the search page. We are an interdisciplinary team with a focus on customer obsession and inventing and simplifying. Our primary focus is on improving the SP experience in search by gaining a deep understanding of shopper pain points and developing new innovative solutions to address them. You will be on the Interleaving team. Our mission is to personalize and contextualize SP ad allocation on the search page. We do this by modeling shopper responses to the number, placement, and quality of ads. We are a data- and hypothesis-driven organization that uses online experimentation, simulation, causal modeling, and online feedback to place ads where they’re useful to shoppers and provide improved discoverability and sales for advertisers. You’ll lead experimentation efforts based on data deep dives to support these solutions. This is a unique opportunity for someone who wants to have broad business impact, a direct impact on customers and the search experience, and get broad exposure to a wide range of scientific techniques (machine learning, bandit learning, optimization, LLMs). We are looking for an Applied Scientist to join Interleaving team in Marketplace Intelligence with a broad mandate to experiment and innovate to grow Sponsored Products. If you thrive in a product-focussed and data-driven environment, then this role is for you. As an Applied Scientist in this team, you will help to identify unique opportunities to create customized and delightful shopping experience for our growing marketplaces worldwide. Your job will be to identify big opportunities for the team that can help to grow Sponsored Products business working with retail partner teams, product managers, software engineers and TPMs. You will have opportunity to design, run and analyze / experiments to improve the experience of millions of Amazon shoppers while driving quantifiable revenue impact. More importantly, you will have the opportunity to broaden your technical skills in an environment that thrives on creativity, experimentation, and product innovation. Key job responsibilities * Tackle and solve challenging science and business problems that balance the interests of advertisers, shoppers, and Amazon. * Develop real-time machine learning algorithms to allocate billions of ads per day in advertising auctions. * Develop efficient algorithms for multi-objective optimization and AI control methods to find operating points for the ad marketplace then evolve them * Be an expert at designing and implementing solutions that use a range of data science methodologies to automate data analysis or to solve complex business problems. * Perform hands-on analysis and modeling of enormous data sets to develop insights that improve shopper experience, without compromising Ad revenue in addition to designing metrics for complex systems. * Drive end-to-end machine learning projects that have a high degree of ambiguity, scale, complexity. * Run A/B experiments, gather data, and perform statistical analysis.
IN, KA, Bengaluru
The Amazon Artificial Generative Intelligence (AGI) team in India is seeking a talented, self-driven Applied Scientist to work on prototyping, optimizing, and deploying ML algorithms within the realm of Generative AI. Key job responsibilities - Research, experiment and build Proof Of Concepts advancing the state of the art in AI & ML for GenAI. - Collaborate with cross-functional teams to architect and execute technically rigorous AI projects. - Thrive in dynamic environments, adapting quickly to evolving technical requirements and deadlines. - Engage in effective technical communication (written & spoken) with coordination across teams. - Conduct thorough documentation of algorithms, methodologies, and findings for transparency and reproducibility. - Publish research papers in internal and external venues of repute - Support on-call activities for critical issues
CN, 13, Beijing
北京职位 - 如果希望在北京工作,请投递本职位。 毕业时间:2024年10月 - 2025年9月之间毕业的应届毕业生 · 投递须知: 1 填写简历申请时,请把必填和非必填项都填写完整。提交简历之后就无法修改了哦! 2 学校的英文全称请准确填写。中英文对应表,请点击链接查看 https://docs.qq.com/sheet/DVmdaa1BCV0RBbnlR?tab=BB08J2 如果您正在攻读NLP,IR或搜索领域专业的博士或硕士研究生,而且对应用科学家的工作感兴趣。如果您也喜爱深入研究棘手的技术问题并提出解决方案,用成功的产品显著地改善人们的生活。 那么,我们诚挚邀请您加入亚马逊的International Technology搜索团队改善Amazon的产品搜索服务。我们的目标是帮助亚马逊的客户找到他们所需的产品,并发现他们感兴趣的新产品。这会是一份收获满满的工作。您每天的工作都与全球数百万亚马逊客户的体验紧密相关。您将提出和探索NLP和IR领域的创新,基于TB级别的产品和流量数据设计机器学习模型。您将集成这些模型到搜索引擎中为客户提供服务,通过数据,建模和客户反馈来完成闭环。您对模型的选择需要能够平衡业务指标和响应时间的需求。
IL, Tel Aviv
Come build the future of entertainment with us. Are you interested in helping shape the future of movies and television? Do you want to help define the next generation of how and what Amazon customers are watching? Prime Video is a premium streaming service that offers customers a vast collection of TV shows and movies - all with the ease of finding what they love to watch in one place. We offer customers thousands of popular movies and TV shows including Amazon Originals and exclusive licensed content to exciting live sports events. We also offer our members the opportunity to subscribe to add-on channels which they can cancel at anytime and to rent or buy new release movies and TV box sets on the Prime Video Store. Prime Video is a fast-paced, growth business - available in over 240 countries and territories worldwide. The team works in a dynamic environment where innovating on behalf of our customers is at the heart of everything we do. If this sounds exciting to you, please read on. We are looking for an Applied Scientist to embark on our journey to build a Prime Video Sports tech team in Israel from ground up. Our team will focus on developing products to allow for personalizing the customers’ experience and providing them real-time insights and revolutionary experiences using Computer Vision (CV) and Machine Learning (ML). You will get a chance to work on greenfield, cutting-edge and large-scale engineering and science projects, and a rare opportunity to be one of the founders of the Israel Prime Video Sports tech team in Israel. Key job responsibilities We are looking for an Applied Scientist with domain expertise in Computer Vision or Recommendation Systems to lead development of new algorithms and E2E solutions. You will be part of a team of applied scientists and software development engineers responsible for research, design, development and deployment of algorithms into production pipelines. As a technologist, you will also drive publications of original work in top-tier conferences in Computer Vision and Machine Learning. You will be expected to deal with ambiguity! We're looking for someone with outstanding analytical abilities and someone comfortable working with cross-functional teams and systems. You must be a self-starter and be able to learn on the go. About the team In September 2018 Prime Video launched its first full-scale live streaming experience to world-wide Prime customers with NFL Thursday Night Football. That was just the start. Now Amazon has exclusive broadcasting rights to major leagues like NFL Thursday Night Football, Tennis major like Roland-Garros and English Premium League to list few and are broadcasting live events across 30 sports world-wide. Prime Video is expanding not just the breadth of live content that it offers, but the depth of the experience. This is a transformative opportunity, the chance to be at the vanguard of a program that will revolutionize Prime Video, and the live streaming experience of customers everywhere.
US, WA, Bellevue
Amazon is looking for world class scientists to join its AWS Fundamental Research Team working within a variety of machine learning disciplines. This group is entrusted with developing core machine learning solutions for AWS services. At the AWS Fundamental Research Team you will invent, implement, and deploy state of the art machine learning algorithms and systems. You will build prototypes and explore conceptually large scale ML solutions across different domains and computation platforms. You will interact closely with our customers and with the academic community. You will be at the heart of a growing and exciting focus area for AWS and work with other acclaimed engineers and world famous scientists. This team is part of AWS Utility Computing: Utility Computing (UC) AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry. As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Hybrid Work We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face. Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.
US, WA, Bellevue
Amazon is looking for world class scientists to join its AWS Fundamental Research Team working within a variety of machine learning disciplines. This group is entrusted with developing core machine learning solutions for AWS services. At the AWS Fundamental Research Team you will invent, implement, and deploy state of the art machine learning algorithms and systems. You will build prototypes and explore conceptually large scale ML solutions across different domains and computation platforms. You will interact closely with our customers and with the academic community. You will be at the heart of a growing and exciting focus area for AWS and work with other acclaimed engineers and world famous scientists. This team is part of AWS Utility Computing: Utility Computing (UC) AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry. As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Hybrid Work We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face. Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.
US, WA, Bellevue
Amazon is looking for world class senior applied scientist to join its AWS Fundamental Research Team working within a variety of machine learning disciplines. This group is entrusted with developing core machine learning solutions for AWS services. At the AWS Fundamental Research Team you will invent, implement, and deploy state of the art machine learning algorithms and systems. You will build prototypes and explore conceptually large scale ML solutions across different domains and computation platforms. You will interact closely with our customers and with the academic community. You will be at the heart of a growing and exciting focus area for AWS and work with other acclaimed engineers and world famous scientists. This team is part of AWS Utility Computing: Utility Computing (UC) AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry. As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Hybrid Work We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face. Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.
US, WA, Bellevue
Amazon is looking for world class scientists to join its AWS Fundamental Research Team working within a variety of machine learning disciplines. This group is entrusted with developing core machine learning solutions for AWS services. At the AWS Fundamental Research Team you will invent, implement, and deploy state of the art machine learning algorithms and systems. You will build prototypes and explore conceptually large scale ML solutions across different domains and computation platforms. You will interact closely with our customers and with the academic community. You will be at the heart of a growing and exciting focus area for AWS and work with other acclaimed engineers and world famous scientists. This team is part of AWS Utility Computing: Utility Computing (UC) AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS’s services and features apart in the industry. As a member of the UC organization, you’ll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Hybrid Work We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face. Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.