🔍

AWS Certified Machine Learning - Specialty MLS-C01 — Question 71

Topic 1 · Question 71 of 369

Topic 1 · Question 71

A Data Scientist needs to migrate an existing on-premises ETL process to the cloud. The current process runs at regular time intervals and uses PySpark to combine and format multiple large data sources into a single consolidated output for downstream processing. The Data Scientist has been given the following requirements to the cloud solution: • Combine multiple data sources. • Reuse existing PySpark logic. • Run the solution on the existing schedule. • Minimize the number of servers that will need to be managed. Which architecture should the Data Scientist use to build this solution?