πŸ”

PDE β€” questions

Page 5 of 18 Β· 341 total questions.

Topic 1 Β· Question 81

MJTelco Case Study - Company Overview - MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware. Company Background - Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs. Solution Concept - MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs: β€’ Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations. Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition. MJTelco will also use three separate operating environments `" development/test, staging, and production `" to meet the needs of running experiments, deploying new features, and serving production customers. Business Requirements - β€’ Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. β€’ Ensure security of their proprietary data to protect their leading-edge machine learning and analysis. β€’ Provide reliable and timely access to data for analysis from distributed research workers β€’ Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers. Technical Requirements - Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each. Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles. CEO Statement - Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments. CTO Statement - Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate. CFO Statement - The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines. You need to compose visualization for operations teams with the following requirements: β€’ Telemetry must include data from all 50,000 installations for the most recent 6 weeks (sampling once every minute) β€’ The report must not be more than 3 hours delayed from live data. β€’ The actionable report should only show suboptimal links. β€’ Most suboptimal links should be sorted to the top. Suboptimal links can be grouped and filtered by regional geography. β€’ User response time to load the report must be <5 seconds. You create a data source to store the last 6 weeks of data, and create visualizations that allow viewers to see multiple date ranges, distinct geographic regions, and unique installation types. You always show the latest data without any changes to your visualizations. You want to avoid creating and updating new visualizations each month. What should you do?

  • ALook through the current data and compose a series of charts and tables, one for each possible combination of criteria.
  • BLook through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection. (correct answer)
  • CExport the data to a spreadsheet, compose a series of charts and tables, one for each possible combination of criteria, and spread them across multiple tabs.
  • DLoad the data into relational database tables, write a Google App Engine application that queries all rows, summarizes the data across each criteria, and then renders results using the Google Charts and visualization API.
Reveal answer & explanation
Correct answer: B

The correct answer is B. Option B: Look through the current data and compose a small set of generalized charts and tables bound to criteria filters that allow value selection. This option meets the real-time / low-latency performance requirement.

Topic 1 Β· Question 82

MJTelco Case Study - Company Overview - MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware. Company Background - Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to-many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs. Solution Concept - MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs: β€’ Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations. β€’ Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition. MJTelco will also use three separate operating environments `" development/test, staging, and production `" to meet the needs of running experiments, deploying new features, and serving production customers. Business Requirements - β€’ Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. β€’ Ensure security of their proprietary data to protect their leading-edge machine learning and analysis. β€’ Provide reliable and timely access to data for analysis from distributed research workers β€’ Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers. Technical Requirements - Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each. Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately 100m records/day Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles. CEO Statement - Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments. CTO Statement - Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate. CFO Statement - The project is too large for us to maintain the hardware and software required for the data and analysis. Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines. Given the record streams MJTelco is interested in ingesting per day, they are concerned about the cost of Google BigQuery increasing. MJTelco asks you to provide a design solution. They require a single large data table called tracking_table. Additionally, they want to minimize the cost of daily queries while performing fine-grained analysis of each day's events. They also want to use streaming ingestion. What should you do?

  • ACreate a table called tracking_table and include a DATE column.
  • BCreate a partitioned table called tracking_table and include a TIMESTAMP column. (correct answer)
  • CCreate sharded tables for each day following the pattern tracking_table_YYYYMMDD.
  • DCreate a table called tracking_table with a TIMESTAMP column to represent the day.
Reveal answer & explanation
Correct answer: B

The correct answer is B. Option B: Create a partitioned table called tracking_table and include a TIMESTAMP column. This option meets the real-time / low-latency performance requirement.

Topic 1 Β· Question 83

Flowlogistic Case Study - Company Overview - Flowlogistic is a leading logistics and supply chain provider. They help businesses throughout the world manage their resources and transport them to their final destination. The company has grown rapidly, expanding their offerings to include rail, truck, aircraft, and oceanic shipping. Company Background - The company started as a regional trucking company, and then expanded into other logistics market. Because they have not updated their infrastructure, managing and tracking orders and shipments has become a bottleneck. To improve operations, Flowlogistic developed proprietary technology for tracking shipments in real time at the parcel level. However, they are unable to deploy it because their technology stack, based on Apache Kafka, cannot support the processing volume. In addition, Flowlogistic wants to further analyze their orders and shipments to determine how best to deploy their resources. Solution Concept - Flowlogistic wants to implement two concepts using the cloud: β€’ Use their proprietary technology in a real-time inventory-tracking system that indicates the location of their loads β€’ Perform analytics on all their orders and shipment logs, which contain both structured and unstructured data, to determine how best to deploy resources, which markets to expand info. They also want to use predictive analytics to learn earlier when a shipment will be delayed. Existing Technical Environment - Flowlogistic architecture resides in a single data center: β€’ Databases - 8 physical servers in 2 clusters - SQL Server `" user data, inventory, static data - 3 physical servers - Cassandra `" metadata, tracking messages 10 Kafka servers `" tracking message aggregation and batch insert β€’ Application servers `" customer front end, middleware for order/customs - 60 virtual machines across 20 physical servers - Tomcat `" Java services - Nginx `" static content - Batch servers β€’ Storage appliances - iSCSI for virtual machine (VM) hosts - Fibre Channel storage area network (FC SAN) `" SQL server storage Network-attached storage (NAS) image storage, logs, backups β€’ 10 Apache Hadoop /Spark servers - Core Data Lake - Data analysis workloads β€’ 20 miscellaneous servers - Jenkins, monitoring, bastion hosts, Business Requirements - β€’ Build a reliable and reproducible environment with scaled panty of production. β€’ Aggregate data in a centralized Data Lake for analysis β€’ Use historical data to perform predictive analytics on future shipments β€’ Accurately track every shipment worldwide using proprietary technology β€’ Improve business agility and speed of innovation through rapid provisioning of new resources β€’ Analyze and optimize architecture for performance in the cloud β€’ Migrate fully to the cloud if all other requirements are met Technical Requirements - β€’ Handle both streaming and batch data β€’ Migrate existing Hadoop workloads β€’ Ensure architecture is scalable and elastic to meet the changing demands of the company. β€’ Use managed services whenever possible β€’ Encrypt data flight and at rest Connect a VPN between the production data center and cloud environment SEO Statement - We have grown so quickly that our inability to upgrade our infrastructure is really hampering further growth and efficiency. We are efficient at moving shipments around the world, but we are inefficient at moving data around. We need to organize our information so we can more easily understand where our customers are and what they are shipping. CTO Statement - IT has never been a priority for us, so as our data has grown, we have not invested enough in our technology. I have a good staff to manage IT, but they are so busy managing our infrastructure that I cannot get them to do the things that really matter, such as organizing our data, building the analytics, and figuring out how to implement the CFO' s tracking technology. CFO Statement - Part of our competitive advantage is that we penalize ourselves for late shipments and deliveries. Knowing where out shipments are at all times has a direct correlation to our bottom line and profitability. Additionally, I don't want to commit capital to building out a server environment. Flowlogistic's management has determined that the current Apache Kafka servers cannot handle the data volume for their real-time inventory tracking system. You need to build a new system on Google Cloud Platform (GCP) that will feed the proprietary tracking software. The system must be able to ingest data from a variety of global sources, process and query in real-time, and store the data reliably. Which combination of GCP products should you choose?

  • ACloud Pub/Sub, Cloud Dataflow, and Cloud Storage (correct answer)
  • BCloud Pub/Sub, Cloud Dataflow, and Local SSD
  • CCloud Pub/Sub, Cloud SQL, and Cloud Storage
  • DCloud Load Balancing, Cloud Dataflow, and Cloud Storage
  • ECloud Dataflow, Cloud SQL, and Cloud Storage
Reveal answer & explanation
Correct answer: A

The correct answer is A. Option A: Cloud Pub/Sub, Cloud Dataflow, and Cloud Storage

Explanation

Cloud Storage provides durable, scalable object storage that is fully managed. Dataflow runs serverless Apache Beam pipelines for stream and batch data processing with autoscaling. Pub/Sub is a serverless, global messaging service that decouples services and ingests high-volume event streams. This option meets the real-time / low-latency performance requirement.

Topic 1 Β· Question 84

After migrating ETL jobs to run on BigQuery, you need to verify that the output of the migrated jobs is the same as the output of the original. You've loaded a table containing the output of the original job and want to compare the contents with output from the migrated job to show that they are identical. The tables do not contain a primary key column that would enable you to join them together for comparison. What should you do?

  • ASelect random samples from the tables using the RAND() function and compare the samples.
  • BSelect random samples from the tables using the HASH() function and compare the samples.
  • CUse a Dataproc cluster and the BigQuery Hadoop connector to read the data from each table and calculate a hash from non-timestamp columns of the table after sorting. Compare the hashes of each table. (correct answer)
  • DCreate stratified random samples using the OVER() function and compare equivalent samples from each table.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Use a Dataproc cluster and the BigQuery Hadoop connector to read the data from each table and calculate a hash from non-timestamp columns of the table after sorting. Compare the hashes of each table.

Explanation

BigQuery is a serverless, petabyte-scale data warehouse for fast SQL analytics with no infrastructure to manage. Dataproc runs managed Spark and Hadoop clusters for big-data processing.

Topic 1 Β· Question 85

You are a head of BI at a large enterprise company with multiple business units that each have different priorities and budgets. You use on-demand pricing for BigQuery with a quota of 2K concurrent on-demand slots per project. Users at your organization sometimes don't get slots to execute their query and you need to correct this. You'd like to avoid introducing new projects to your account. What should you do?

  • AConvert your batch BQ queries into interactive BQ queries.
  • BCreate an additional project to overcome the 2K on-demand per-project quota.
  • CSwitch to flat-rate pricing and establish a hierarchical priority model for your projects. (correct answer)
  • DIncrease the amount of concurrent slots per project at the Quotas page at the Cloud Console.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Switch to flat-rate pricing and establish a hierarchical priority model for your projects.

Topic 1 Β· Question 86

You have an Apache Kafka cluster on-prem with topics containing web application logs. You need to replicate the data to Google Cloud for analysis in BigQuery and Cloud Storage. The preferred replication method is mirroring to avoid deployment of Kafka Connect plugins. What should you do?

  • ADeploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS. (correct answer)
  • BDeploy a Kafka cluster on GCE VM Instances with the Pub/Sub Kafka connector configured as a Sink connector. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
  • CDeploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Source connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
  • DDeploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Sink connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
Reveal answer & explanation
Correct answer: A

The correct answer is A. Option A: Deploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.

Explanation

Dataflow runs serverless Apache Beam pipelines for stream and batch data processing with autoscaling. Dataproc runs managed Spark and Hadoop clusters for big-data processing.

Topic 1 Β· Question 87

You've migrated a Hadoop job from an on-prem cluster to dataproc and GCS. Your Spark job is a complicated analytical workload that consists of many shuffling operations and initial data are parquet files (on average 200-400 MB size each). You see some degradation in performance after the migration to Dataproc, so you'd like to optimize for it. You need to keep in mind that your organization is very cost-sensitive, so you'd like to continue using Dataproc on preemptibles (with 2 non-preemptible workers only) for this workload. What should you do?

  • AIncrease the size of your parquet files to ensure them to be 1 GB minimum. (correct answer)
  • BSwitch to TFRecords formats (appr. 200MB per file) instead of parquet files.
  • CSwitch from HDDs to SSDs, copy initial data from GCS to HDFS, run the Spark job and copy results back to GCS.
  • DSwitch from HDDs to SSDs, override the preemptible VMs configuration to increase the boot disk size.
Reveal answer & explanation
Correct answer: A

The correct answer is A. Option A: Increase the size of your parquet files to ensure them to be 1 GB minimum.

Topic 1 Β· Question 88

Your team is responsible for developing and maintaining ETLs in your company. One of your Dataflow jobs is failing because of some errors in the input data, and you need to improve reliability of the pipeline (incl. being able to reprocess all failing data). What should you do?

  • AAdd a filtering step to skip these types of errors in the future, extract erroneous rows from logs.
  • BAdd a try… catch block to your DoFn that transforms the data, extract erroneous rows from logs.
  • CAdd a try… catch block to your DoFn that transforms the data, write erroneous rows to Pub/Sub PubSub directly from the DoFn.
  • DAdd a try… catch block to your DoFn that transforms the data, use a sideOutput to create a PCollection that can be stored to Pub/Sub later. (correct answer)
Reveal answer & explanation
Correct answer: D

The correct answer is D. Option D: Add a try… catch block to your DoFn that transforms the data, use a sideOutput to create a PCollection that can be stored to Pub/Sub later.

Explanation

Pub/Sub is a serverless, global messaging service that decouples services and ingests high-volume event streams.

Topic 1 Β· Question 89

You're training a model to predict housing prices based on an available dataset with real estate properties. Your plan is to train a fully connected neural net, and you've discovered that the dataset contains latitude and longitude of the property. Real estate professionals have told you that the location of the property is highly influential on price, so you'd like to engineer a feature that incorporates this physical dependency. What should you do?

  • AProvide latitude and longitude as input vectors to your neural net.
  • BCreate a numeric column from a feature cross of latitude and longitude.
  • CCreate a feature cross of latitude and longitude, bucketize it at the minute level and use L1 regularization during optimization. (correct answer)
  • DCreate a feature cross of latitude and longitude, bucketize it at the minute level and use L2 regularization during optimization.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Create a feature cross of latitude and longitude, bucketize it at the minute level and use L1 regularization during optimization.

Topic 1 Β· Question 90

You are deploying MariaDB SQL databases on GCE VM Instances and need to configure monitoring and alerting. You want to collect metrics including network connections, disk IO and replication status from MariaDB with minimal development effort and use StackDriver for dashboards and alerts. What should you do?

  • AInstall the OpenCensus Agent and create a custom metric collection application with a StackDriver exporter.
  • BPlace the MariaDB instances in an Instance Group with a Health Check.
  • CInstall the StackDriver Logging Agent and configure fluentd in_tail plugin to read MariaDB logs.
  • DInstall the StackDriver Agent and configure the MySQL plugin. (correct answer)
Reveal answer & explanation
Correct answer: D

The correct answer is D. Option D: Install the StackDriver Agent and configure the MySQL plugin.

Explanation

Cloud Operations (formerly Stackdriver) provides monitoring, logging, and tracing for reliability.

Topic 1 Β· Question 91

You work for a bank. You have a labelled dataset that contains information on already granted loan application and whether these applications have been defaulted. You have been asked to train a model to predict default rates for credit applicants. What should you do?

  • AIncrease the size of the dataset by collecting additional data.
  • BTrain a linear regression to predict a credit default risk score. (correct answer)
  • CRemove the bias from the data and collect applications that have been declined loans.
  • DMatch loan applicants with their social profiles to enable feature engineering.
Reveal answer & explanation
Correct answer: B

The correct answer is B. Option B: Train a linear regression to predict a credit default risk score.

Topic 1 Β· Question 92

You need to migrate a 2TB relational database to Google Cloud Platform. You do not have the resources to significantly refactor the application that uses this database and cost to operate is of primary concern. Which service do you select for storing and serving your data?

  • ACloud Spanner
  • BCloud Bigtable
  • CCloud Firestore
  • DCloud SQL (correct answer)
Reveal answer & explanation
Correct answer: D

The correct answer is D. Option D: Cloud SQL

Explanation

Cloud SQL is a managed relational database (MySQL/PostgreSQL/SQL Server) that handles patching, backups, and failover.

Topic 1 Β· Question 93

You're using Bigtable for a real-time application, and you have a heavy load that is a mix of read and writes. You've recently identified an additional use case and need to perform hourly an analytical job to calculate certain statistics across the whole database. You need to ensure both the reliability of your production application as well as the analytical workload. What should you do?

  • AExport Bigtable dump to GCS and run your analytical job on top of the exported files.
  • BAdd a second cluster to an existing instance with a multi-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.
  • CAdd a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload. (correct answer)
  • DIncrease the size of your existing cluster twice and execute your analytics workload on your new resized cluster.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Add a second cluster to an existing instance with a single-cluster routing, use live-traffic app profile for your regular workload and batch-analytics profile for the analytics workload.

Explanation

Google Cloud Batch schedules and runs batch jobs at scale without managing infrastructure. This option meets the real-time / low-latency performance requirement.

Topic 1 Β· Question 94

You are designing an Apache Beam pipeline to enrich data from Cloud Pub/Sub with static reference data from BigQuery. The reference data is small enough to fit in memory on a single worker. The pipeline should write enriched results to BigQuery for analysis. Which job type and transforms should this pipeline use?

  • ABatch job, PubSubIO, side-inputs
  • BStreaming job, PubSubIO, JdbcIO, side-outputs
  • CStreaming job, PubSubIO, BigQueryIO, side-inputs (correct answer)
  • DStreaming job, PubSubIO, BigQueryIO, side-outputs
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Streaming job, PubSubIO, BigQueryIO, side-inputs

Explanation

BigQuery is a serverless, petabyte-scale data warehouse for fast SQL analytics with no infrastructure to manage.

Topic 1 Β· Question 95 Β· Select all that apply

You have a data pipeline that writes data to Cloud Bigtable using well-designed row keys. You want to monitor your pipeline to determine when to increase the size of your Cloud Bigtable cluster. Which two actions can you take to accomplish this? (Choose two.)

  • AReview Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Read pressure index is above 100.
  • BReview Key Visualizer metrics. Increase the size of the Cloud Bigtable cluster when the Write pressure index is above 100.
  • CMonitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency. (correct answer)
  • DMonitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity. (correct answer)
  • EMonitor latency of read operations. Increase the size of the Cloud Bigtable cluster of read operations take longer than 100 ms.
Reveal answer & explanation
Correct answer: C, D

The correct answer is C, D. Option C: Monitor the latency of write operations. Increase the size of the Cloud Bigtable cluster when there is a sustained increase in write latency. Option D: Monitor storage utilization. Increase the size of the Cloud Bigtable cluster when utilization increases above 70% of max capacity.

Explanation

Cloud Bigtable is a managed, low-latency NoSQL wide-column store for very high-throughput workloads.

Topic 1 Β· Question 96

You want to analyze hundreds of thousands of social media posts daily at the lowest cost and with the fewest steps. You have the following requirements: β€’ You will batch-load the posts once per day and run them through the Cloud Natural Language API. β€’ You will extract topics and sentiment from the posts. β€’ You must store the raw posts for archiving and reprocessing. β€’ You will create dashboards to be shared with people both inside and outside your organization. You need to store both the data extracted from the API to perform analysis as well as the raw social media posts for historical archiving. What should you do?

  • AStore the social media posts and the data extracted from the API in BigQuery.
  • BStore the social media posts and the data extracted from the API in Cloud SQL.
  • CStore the raw social media posts in Cloud Storage, and write the data extracted from the API into BigQuery. (correct answer)
  • DFeed to social media posts into the API directly from the source, and write the extracted data from the API into BigQuery.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Store the raw social media posts in Cloud Storage, and write the data extracted from the API into BigQuery.

Explanation

Cloud Storage provides durable, scalable object storage that is fully managed. BigQuery is a serverless, petabyte-scale data warehouse for fast SQL analytics with no infrastructure to manage. This option delivers the requirement at the lowest cost.

Topic 1 Β· Question 97

You store historic data in Cloud Storage. You need to perform analytics on the historic data. You want to use a solution to detect invalid data entries and perform data transformations that will not require programming or knowledge of SQL. What should you do?

  • AUse Cloud Dataflow with Beam to detect errors and perform transformations.
  • BUse Cloud Dataprep with recipes to detect errors and perform transformations. (correct answer)
  • CUse Cloud Dataproc with a Hadoop job to detect errors and perform transformations.
  • DUse federated tables in BigQuery with queries to detect errors and perform transformations.
Reveal answer & explanation
Correct answer: B

The correct answer is B. Option B: Use Cloud Dataprep with recipes to detect errors and perform transformations.

Explanation

Dataprep visually explores, cleans, and prepares data for analysis with no code.

Topic 1 Β· Question 98

Your company needs to upload their historic data to Cloud Storage. The security rules don't allow access from external IPs to their on-premises resources. After an initial upload, they will add new data from existing on-premises applications every day. What should they do?

  • AExecute gsutil rsync from the on-premises servers. (correct answer)
  • BUse Dataflow and write the data to Cloud Storage.
  • CWrite a job template in Dataproc to perform the data transfer.
  • DInstall an FTP server on a Compute Engine VM to receive the files and move them to Cloud Storage.
Reveal answer & explanation
Correct answer: A

The correct answer is A. Option A: Execute gsutil rsync from the on-premises servers.

Topic 1 Β· Question 99

You have a query that filters a BigQuery table using a WHERE clause on timestamp and ID columns. By using bq query `"-dry_run you learn that the query triggers a full scan of the table, even though the filter on timestamp and ID select a tiny fraction of the overall data. You want to reduce the amount of data scanned by BigQuery with minimal changes to existing SQL queries. What should you do?

  • ACreate a separate table for each ID.
  • BUse the LIMIT keyword to reduce the number of rows returned.
  • CRecreate the table with a partitioning column and clustering column. (correct answer)
  • DUse the bq query --maximum_bytes_billed flag to restrict the number of bytes billed.
Reveal answer & explanation
Correct answer: C

The correct answer is C. Option C: Recreate the table with a partitioning column and clustering column.

Topic 1 Β· Question 100

You have a requirement to insert minute-resolution data from 50,000 sensors into a BigQuery table. You expect significant growth in data volume and need the data to be available within 1 minute of ingestion for real-time analysis of aggregated trends. What should you do?

  • AUse bq load to load a batch of sensor data every 60 seconds.
  • BUse a Cloud Dataflow pipeline to stream data into the BigQuery table. (correct answer)
  • CUse the INSERT statement to insert a batch of data every 60 seconds.
  • DUse the MERGE statement to apply updates in batch every 60 seconds.
Reveal answer & explanation
Correct answer: B

The correct answer is B. Option B: Use a Cloud Dataflow pipeline to stream data into the BigQuery table.

Explanation

BigQuery is a serverless, petabyte-scale data warehouse for fast SQL analytics with no infrastructure to manage. Dataflow runs serverless Apache Beam pipelines for stream and batch data processing with autoscaling. This option meets the real-time / low-latency performance requirement.

Showing questions 81–100 of 341 Β· Page 5 of 18