Correct Answer: Google Cloud's unified machine learning platform
Explanation:
Vertex AI is Google Cloud's unified platform that brings together tools for building, deploying, and scaling machine learning and generative AI models.
Incorrect! Try again.
2Which company provides the Vertex AI platform?
Overview of Vertex AI
Easy
A.Google Cloud
B.Microsoft Azure
C.IBM Cloud
D.Amazon Web Services
Correct Answer: Google Cloud
Explanation:
Vertex AI is a product offered by Google Cloud for developing and managing machine learning workflows.
Incorrect! Try again.
3What is the main purpose of Generative AI Studio?
Generative AI Studio
Easy
A.To prototype and test generative AI prompts and models
B.To store large datasets in a warehouse
C.To manage cloud billing accounts
D.To monitor network traffic security
Correct Answer: To prototype and test generative AI prompts and models
Explanation:
Generative AI Studio provides an interactive environment to design, test, and tune prompts and generative models before deployment.
Incorrect! Try again.
4Which activity is commonly performed in Generative AI Studio?
Generative AI Studio
Easy
A.Firewall configuration
B.Prompt engineering and testing
C.Domain name registration
D.Physical server maintenance
Correct Answer: Prompt engineering and testing
Explanation:
Generative AI Studio is designed for prompt engineering, allowing users to experiment with and refine prompts interactively.
Incorrect! Try again.
5What type of model is Gemini?
Gemini Models
Easy
A.A file compression algorithm
B.A relational database engine
C.A multimodal generative AI model
D.A spreadsheet automation tool
Correct Answer: A multimodal generative AI model
Explanation:
Gemini is Google's family of multimodal generative AI models capable of processing text, images, audio, and more.
Incorrect! Try again.
6The term multimodal in Gemini models means the model can process:
Gemini Models
Easy
A.Multiple types of input such as text, images, and audio
B.Only a single language
C.Only plain text input
D.Only numerical spreadsheet data
Correct Answer: Multiple types of input such as text, images, and audio
Explanation:
Multimodal means the model can understand and work with several data types, including text, images, audio, and video.
Incorrect! Try again.
7What is Model Garden in Vertex AI?
Model Garden
Easy
A.A tool for gardening simulations
B.A billing dashboard for cloud usage
C.A repository of ready-to-use foundation and ML models
D.A network monitoring service
Correct Answer: A repository of ready-to-use foundation and ML models
Explanation:
Model Garden is a curated collection of foundation models and ML models that users can discover, test, and deploy in Vertex AI.
Incorrect! Try again.
8Which task can you perform using Model Garden?
Model Garden
Easy
A.Design database schemas
B.Discover and select pre-built models
C.Encrypt user passwords
D.Configure DNS records
Correct Answer: Discover and select pre-built models
Explanation:
Model Garden lets users browse, discover, and select from a variety of pre-built and foundation models for their applications.
Incorrect! Try again.
9Which is typically the first step in an end-to-end AI development workflow?
End-to-End AI Development Workflow
Easy
A.Real-time monitoring of predictions
B.Model deployment to production
C.Data collection and preparation
D.Retiring the finished model
Correct Answer: Data collection and preparation
Explanation:
AI development usually begins with collecting and preparing data, which forms the foundation for training and evaluating models.
Incorrect! Try again.
10What does model deployment refer to in the AI workflow?
End-to-End AI Development Workflow
Easy
A.Deleting unused training data
B.Choosing a programming language
C.Making a trained model available for use
D.Writing the initial project proposal
Correct Answer: Making a trained model available for use
Explanation:
Deployment is the step where a trained model is made available to serve predictions to users or applications.
Incorrect! Try again.
11What does text-to-image generation do?
Text-to-Image Generation
Easy
A.Compresses image file sizes
B.Converts images into text captions
C.Translates text between languages
D.Creates images from text descriptions
Correct Answer: Creates images from text descriptions
Explanation:
Text-to-image generation produces new images based on a textual prompt describing the desired content.
Incorrect! Try again.
12In text-to-image generation, the prompt is:
Text-to-Image Generation
Easy
A.A database query for images
B.The final rendered image file
C.A text description guiding the image creation
D.A hardware acceleration setting
Correct Answer: A text description guiding the image creation
Explanation:
The prompt is the text input that describes what the model should generate as an image.
Incorrect! Try again.
13What is the goal of image understanding?
Image Understanding
Easy
A.To convert images into audio
B.To generate images from text
C.To analyze and interpret the content of images
D.To increase image resolution only
Correct Answer: To analyze and interpret the content of images
Explanation:
Image understanding involves analyzing images to identify objects, scenes, text, or other meaningful content.
Incorrect! Try again.
14Which task is an example of image understanding?
Image Understanding
Easy
A.Encrypting an image for storage
B.Renaming an image file
C.Adjusting screen brightness
D.Identifying objects present in a photo
Correct Answer: Identifying objects present in a photo
Explanation:
Detecting and labeling objects within an image is a core image understanding task.
Incorrect! Try again.
15What does video analysis typically involve?
Video Analysis
Easy
A.Playing videos at higher speed
B.Reducing video download times
C.Extracting information from video content
D.Adding subtitles manually
Correct Answer: Extracting information from video content
Explanation:
Video analysis uses AI to extract insights such as objects, actions, or scenes from video content.
Incorrect! Try again.
16Which of the following is a common video analysis task?
Video Analysis
Easy
A.Detecting actions or events in a video
B.Compressing an audio file
C.Sorting emails by date
D.Formatting a text document
Correct Answer: Detecting actions or events in a video
Explanation:
Recognizing actions, objects, or events over time is a typical video analysis task.
Incorrect! Try again.
17Which is an example of an audio processing task?
Audio Processing
Easy
A.Cropping an image
B.Rendering a 3D model
C.Indexing a database table
D.Transcribing speech into text
Correct Answer: Transcribing speech into text
Explanation:
Speech-to-text transcription is a common audio processing task where spoken audio is converted into written text.
Incorrect! Try again.
18Speech-to-text technology converts:
Audio Processing
Easy
A.Spoken words into written text
B.Text into printed documents
C.Images into descriptions
D.Video into slideshows
Correct Answer: Spoken words into written text
Explanation:
Speech-to-text takes audio of spoken language and produces the corresponding written text.
Incorrect! Try again.
19What is cross-modal reasoning?
Cross-Modal Reasoning
Easy
A.Reasoning using only text data
B.Reasoning across different data types like text and images
C.Sorting files by size
D.Encrypting data across servers
Correct Answer: Reasoning across different data types like text and images
Explanation:
Cross-modal reasoning combines information from multiple modalities, such as text and images, to reach conclusions.
Incorrect! Try again.
20A multimodal application is one that:
Multimodal Application Design
Easy
A.Works with multiple types of data such as text, images, and audio
B.Only processes plain text files
C.Runs exclusively on mobile devices
D.Handles a single spreadsheet at a time
Correct Answer: Works with multiple types of data such as text, images, and audio
Explanation:
Multimodal applications are designed to handle and combine several data types, like text, images, and audio, within a single system.
Incorrect! Try again.
21A data science team wants a single managed platform to train custom models, deploy them to endpoints, and also access pre-trained foundation models without stitching together separate services. Which characteristic of Vertex AI best addresses this need?
Overview of Vertex AI
Medium
A.It is a database service optimized for storing model weights
B.It is a unified ML platform that combines data prep, training, deployment, and foundation model access
C.It is a standalone GPU rental service with no MLOps tooling
D.It is a code editor for writing Python training scripts only
Correct Answer: It is a unified ML platform that combines data prep, training, deployment, and foundation model access
Explanation:
Vertex AI is Google Cloud's unified ML platform that brings together data preparation, custom training, deployment/serving, MLOps, and access to foundation models in one environment, avoiding the need to integrate disparate tools.
Incorrect! Try again.
22A developer wants to quickly experiment with prompts, tune generation parameters like temperature, and test model responses through a visual interface before writing production code. Which tool is most appropriate?
Generative AI Studio
Medium
A.Dataproc
B.Generative AI Studio
C.BigQuery
D.Cloud Storage
Correct Answer: Generative AI Studio
Explanation:
Generative AI Studio provides a UI for rapid prompt experimentation, parameter tuning (e.g., temperature, top-k), and testing model outputs, making it ideal for prototyping before coding.
Incorrect! Try again.
23Gemini models are frequently described as natively multimodal. What does this most directly imply for an application developer?
Gemini Models
Medium
A.Multimodality requires converting all inputs to text first
B.The model can only handle text but generates images via plugins
C.A single model can process text, images, audio, and video inputs together
D.Separate specialized models must be called for each data type
Correct Answer: A single model can process text, images, audio, and video inputs together
Explanation:
Natively multimodal means Gemini was designed from the ground up to accept and reason over multiple modalities (text, image, audio, video) within one model, rather than bolting together separate single-modality models.
Incorrect! Try again.
24A team is evaluating whether to use a Google first-party model, an open-source model, or a third-party partner model for their project. Which Vertex AI feature lets them discover, compare, and deploy from all these categories in one place?
Model Garden
Medium
A.Model Garden
B.Cloud Monitoring
C.Artifact Registry
D.Generative AI Studio
Correct Answer: Model Garden
Explanation:
Model Garden is a curated repository within Vertex AI where users can discover, test, and deploy first-party, open-source, and third-party models from a single catalog.
Incorrect! Try again.
25In a typical end-to-end AI development workflow on Vertex AI, which sequence correctly orders the core phases?
B.Model training → data preparation → deployment → evaluation → monitoring
C.Deployment → data preparation → monitoring → training → evaluation
D.Evaluation → deployment → data preparation → training → monitoring
Correct Answer: Data preparation → model training/tuning → evaluation → deployment → monitoring
Explanation:
The standard ML lifecycle starts with preparing data, then training or tuning the model, evaluating performance, deploying to an endpoint, and finally monitoring for drift and quality in production.
Incorrect! Try again.
26When using a text-to-image model, a user's prompt produces images that ignore a key requested detail. Which prompting adjustment most directly improves adherence to that detail?
Text-to-Image Generation
Medium
A.Reducing the prompt to a single vague word
B.Switching the output resolution to the lowest setting
C.Adding specific descriptive keywords and emphasis for the missing detail
D.Removing all adjectives from the prompt
Correct Answer: Adding specific descriptive keywords and emphasis for the missing detail
Explanation:
Text-to-image models respond to prompt specificity; explicitly describing and emphasizing the desired detail increases the likelihood the model renders it, whereas vaguer prompts reduce control.
Incorrect! Try again.
27A retail app sends a product photo to a multimodal model and asks, "How many items are on the shelf and what colors are they?" This task is an example of which capability?
Image Understanding
Medium
A.Audio transcription
B.Image understanding
C.Text-to-image generation
D.Video summarization
Correct Answer: Image understanding
Explanation:
Extracting information such as object counts and attributes (colors) from an input image is image understanding — the model interprets visual content rather than generating an image.
Incorrect! Try again.
28A model is asked to identify the moment a goal is scored in a soccer clip and describe what happened just before it. Which capability does this require beyond simple frame classification?
Video Analysis
Medium
A.Text-to-image rendering of the scene
B.Temporal reasoning over the sequence of frames
C.Static image captioning of a single frame
D.Audio-only keyword spotting
Correct Answer: Temporal reasoning over the sequence of frames
Explanation:
Locating an event and describing preceding actions requires understanding how frames relate over time (temporal reasoning), not just classifying an isolated frame.
Incorrect! Try again.
29A voice-note app needs to convert spoken meeting recordings into searchable text and then summarize action items. Which pipeline correctly reflects the audio processing steps involved?
Audio Processing
Medium
A.Speech-to-text transcription followed by text summarization
B.Text-to-speech synthesis followed by image generation
C.Video analysis followed by text-to-image generation
D.Image understanding followed by audio playback
Correct Answer: Speech-to-text transcription followed by text summarization
Explanation:
First the audio is transcribed to text (speech-to-text), then the resulting text is summarized to extract action items. The other options mismatch the modalities and goal.
Incorrect! Try again.
30An app receives an image of a recipe card and an audio question asking, "Can I make this if I'm allergic to nuts?" Answering requires combining information from both inputs. What is this ability called?
Cross-Modal Reasoning
Medium
A.Cross-modal reasoning
B.Single-modality classification
C.Data augmentation
D.Prompt caching
Correct Answer: Cross-modal reasoning
Explanation:
Cross-modal reasoning is the ability to integrate and reason jointly across different input modalities (here image plus audio) to produce a coherent answer.
Incorrect! Try again.
31When designing a multimodal customer support assistant that accepts screenshots and typed questions, which design decision best improves reliability of responses?
Multimodal Application Design
Medium
A.Randomizing the order of modalities on every request
B.Sending only the image and ignoring the typed question
C.Providing clear instructions and context in the prompt about how to interpret each modality
D.Compressing images to unreadable resolution to save cost
Correct Answer: Providing clear instructions and context in the prompt about how to interpret each modality
Explanation:
Well-structured prompts that tell the model how to use each modality and what the task is lead to more reliable, grounded outputs in multimodal applications.
Incorrect! Try again.
32A developer needs the lowest-latency, cost-efficient Gemini variant for high-volume, simple classification tasks, while another team needs maximum reasoning quality for complex analysis. What is the key trade-off guiding this choice?
Gemini Models
Medium
A.Balancing model size/capability against latency and cost
B.Selecting a model based solely on its release date
C.Picking the model with the largest training dataset regardless of task
D.Choosing between text-only and image-only models
Correct Answer: Balancing model size/capability against latency and cost
Explanation:
Gemini offers variants tuned for different points on the capability-vs-efficiency spectrum; lighter models favor speed and cost, while larger models favor deeper reasoning. The right choice depends on task complexity and volume.
Incorrect! Try again.
33A team wants to reduce hallucinations by having their Vertex AI model reference their internal document store when answering. Which approach fits within the Vertex AI ecosystem?
Overview of Vertex AI
Medium
A.Disabling all safety filters
B.Grounding the model with retrieval from their data (RAG)
C.Reducing the prompt to a single word
D.Increasing the temperature parameter to maximum
Correct Answer: Grounding the model with retrieval from their data (RAG)
Explanation:
Grounding via retrieval-augmented generation lets the model cite and use the team's own documents, improving factual accuracy and reducing hallucinations — a supported pattern in Vertex AI.
Incorrect! Try again.
34In Generative AI Studio, a developer lowers the temperature setting close to 0 for a factual Q&A assistant. What effect should they expect?
Generative AI Studio
Medium
A.Longer responses regardless of prompt
B.More deterministic, focused outputs with less randomness
C.More creative and varied outputs with higher randomness
D.Automatic switching to an image model
Correct Answer: More deterministic, focused outputs with less randomness
Explanation:
Lower temperature reduces sampling randomness, producing more deterministic and consistent responses — desirable for factual tasks. Higher temperature increases creativity and variability.
Incorrect! Try again.
35A designer wants generated images to avoid including any text or watermarks. Which prompting technique is most directly suited to this goal?
Text-to-Image Generation
Medium
A.Lowering the image resolution
B.Adding more positive adjectives about text
C.Increasing the number of output tokens
D.Using negative prompts to specify what to exclude
Correct Answer: Using negative prompts to specify what to exclude
Explanation:
Negative prompting explicitly tells the model which elements to avoid (e.g., text, watermarks), giving finer control over unwanted content in the generated image.
Incorrect! Try again.
36A startup wants to fine-tune an open-source model they found in Model Garden on their proprietary dataset before deploying. Which statement about this workflow is accurate?
Model Garden
Medium
A.Open-source models in Model Garden cannot be customized in any way
B.Fine-tuning must be done outside Google Cloud entirely
C.Model Garden only allows viewing model cards, not any deployment
D.Model Garden lets them select the model, then tune and deploy it through Vertex AI tooling
Correct Answer: Model Garden lets them select the model, then tune and deploy it through Vertex AI tooling
Explanation:
Model Garden integrates with Vertex AI so that eligible models can be selected, fine-tuned on custom data, and deployed to endpoints within the same platform.
Incorrect! Try again.
37After deploying a model, a team notices prediction quality degrading over weeks as real-world data shifts. Which workflow phase addresses this, and what is the phenomenon called?
Continuous monitoring in production catches drift — when incoming data distribution diverges from training data — signaling the need for retraining or updates.
Incorrect! Try again.
38A document-processing app uses a multimodal model to read scanned invoices and extract the total amount and vendor name into structured fields. What is the best description of this task?
Image Understanding
Medium
A.Visual information extraction from images into structured data
B.Converting the invoice audio to text
C.Generating a new invoice image from text
D.Rendering the invoice as a 3D model
Correct Answer: Visual information extraction from images into structured data
Explanation:
Reading text and fields from an image and outputting structured values is image understanding applied to information extraction, common in document AI use cases.
Incorrect! Try again.
39A content team wants automatic chapter markers and a summary for hour-long tutorial videos. Which combination of capabilities does the model need?
Video Analysis
Medium
A.Text-to-image generation of thumbnails only
B.Single-frame image captioning only
C.Audio pitch detection only
D.Temporal segmentation plus content summarization across the video
Correct Answer: Temporal segmentation plus content summarization across the video
Explanation:
Generating chapters requires identifying topic boundaries over time (temporal segmentation) and then summarizing each segment, both of which are video analysis capabilities.
Incorrect! Try again.
40A multimodal app must handle cases where a user uploads an image but sends no text, or sends text with no image. Which design principle best manages this variability?
Multimodal Application Design
Medium
A.Always assuming both modalities are present
B.Rejecting all requests that are not both text and image
C.Gracefully handling missing or optional modalities with fallback logic
D.Randomly discarding one modality per request
Correct Answer: Gracefully handling missing or optional modalities with fallback logic
Explanation:
Robust multimodal apps anticipate that inputs may vary and include fallback handling so the system responds sensibly whether one or multiple modalities are provided.
Incorrect! Try again.
41A team migrates a custom training job from a raw Compute Engine setup to Vertex AI to gain managed MLOps capabilities. They require reproducible pipelines, metadata lineage tracking, and automated model registration on successful evaluation. Which combination of Vertex AI components most directly satisfies all three requirements?
Overview of Vertex AI
Hard
A.Vertex AI Workbench, Cloud Logging, and BigQuery ML
B.Vertex AI Feature Store, Cloud Scheduler, and Dataflow
C.Generative AI Studio, Cloud Build, and Artifact Registry
D.Vertex AI Pipelines, Vertex ML Metadata, and Vertex AI Model Registry
Correct Answer: Vertex AI Pipelines, Vertex ML Metadata, and Vertex AI Model Registry
Explanation:
Vertex AI Pipelines orchestrates reproducible workflows, Vertex ML Metadata captures lineage of artifacts and executions, and the Model Registry handles versioned model registration. The other options mix components that do not directly address all three needs.
Incorrect! Try again.
42A prompt engineer notices that identical prompts in Generative AI Studio produce highly variable outputs across runs, harming reproducibility for a compliance report. Which parameter adjustment most directly reduces this stochasticity while preserving the model?
Generative AI Studio
Hard
A.Raise the maxOutputTokens limit to allow longer completions
B.Increase top_p toward to widen the sampling pool
C.Lower the temperature toward (and optionally set top_k=1)
D.Increase the candidateCount so more responses are returned
Correct Answer: Lower the temperature toward (and optionally set top_k=1)
Explanation:
Temperature scales the sharpness of the output distribution; near the model becomes near-deterministic (greedy). Increasing top_p widens sampling (more variance), token limits affect length not randomness, and candidateCount produces more diverse samples.
Incorrect! Try again.
43An application must reason jointly over a -hour video, its audio track, and a set of PDF technical diagrams within a single request. Which capability of Gemini models is the primary enabler, and what is the key practical constraint?
Gemini Models
Hard
A.Fine-tuning on the specific media; the model must be retrained before every multimodal request
B.Automatic external tool-calling; each modality must first be converted to plain text via separate APIs
C.Native multimodal input with a large context window; the total tokens (including media) must fit within the model's context limit
D.Retrieval-augmented generation; all media must be embedded and stored in a vector database beforehand
Correct Answer: Native multimodal input with a large context window; the total tokens (including media) must fit within the model's context limit
Explanation:
Gemini natively ingests mixed modalities in one prompt, but images, audio, and video consume tokens; the combined token count must stay within the context window. The other options describe unnecessary or incorrect prerequisites.
Incorrect! Try again.
44A developer wants to deploy an open-weights model (e.g., a Llama variant) from Model Garden for low-latency, dedicated serving with full control over the endpoint. Which deployment path is most appropriate?
Model Garden
Hard
A.Deploy the model to a Vertex AI Endpoint with provisioned dedicated resources
B.Call it exclusively through the shared serverless Gemini API
C.Access it only via the Generative AI Studio browser playground
D.Export it to BigQuery ML and query it with SQL for each request
Correct Answer: Deploy the model to a Vertex AI Endpoint with provisioned dedicated resources
Explanation:
Dedicated, low-latency serving with endpoint control requires deploying to a Vertex AI Endpoint with provisioned compute. The Gemini API is for Google's proprietary models, BigQuery ML is for SQL-based inference of specific model types, and the playground is for experimentation only.
Incorrect! Try again.
45In a Vertex AI pipeline, a model passes offline evaluation but degrades sharply after deployment as real traffic shifts. Which monitoring signal most directly identifies this specific failure mode, and what is the correct corrective step?
End-to-End AI Development Workflow
Hard
A.Endpoint CPU utilization alerts triggering a node autoscaling change
B.Pipeline DAG compilation errors triggering a code rollback
C.Model Registry version mismatch triggering a re-registration
D.Feature/prediction drift detection triggering a retraining pipeline
Correct Answer: Feature/prediction drift detection triggering a retraining pipeline
Explanation:
A shift between training and live data distributions is drift; Vertex AI Model Monitoring detects feature/prediction drift and can trigger retraining. Resource utilization, compilation errors, and registry mismatches are unrelated to distributional shift.
Incorrect! Try again.
46A designer using a text-to-image model finds that adding a negative prompt like blurry, extra fingers improves quality. Mechanistically, in a diffusion model using classifier-free guidance, what does a negative prompt influence?
Text-to-Image Generation
Hard
A.It changes the random seed so a completely different latent is sampled
B.It reduces the number of denoising steps required to reach the final image
C.It is appended to the positive prompt and increases the total token weight equally
D.It defines the unconditional/undesired direction so guidance steers the denoising away from those concepts
Correct Answer: It defines the unconditional/undesired direction so guidance steers the denoising away from those concepts
Explanation:
In classifier-free guidance the model interpolates between conditional and unconditional predictions; a negative prompt supplies a concept to steer away from during denoising. It does not merely concatenate tokens, alter step count, or reseed the latent.
Incorrect! Try again.
47A multimodal model is asked to count the number of red cars in a dense parking-lot photo and returns an incorrect count. Which prompting or design strategy is most likely to improve counting accuracy without changing the model?
Image Understanding
Hard
A.Ask the model to enumerate and describe each detected car sequentially before giving a total
B.Increase the temperature so the model explores more counting hypotheses
C.Request the answer as a single number with no reasoning to reduce noise
D.Downscale the image aggressively to speed up feature extraction
Correct Answer: Ask the model to enumerate and describe each detected car sequentially before giving a total
Explanation:
Structured, step-by-step enumeration reduces holistic estimation errors on counting tasks. Higher temperature adds noise, suppressing reasoning removes the intermediate structure that aids counting, and aggressive downscaling loses the detail needed to distinguish objects.
Incorrect! Try again.
48When sending a long video to a Gemini model, a developer observes that fine-grained fast motion events are sometimes missed. Which factor most directly explains this and what is the appropriate mitigation?
Video Analysis
Hard
A.The audio track is overriding the video; mute the audio before upload
B.The context window is too large; reduce it to force denser attention
C.The resolution is too high; downscale frames to improve motion capture
D.The frame sampling rate is too low; increase sampling density or segment the video into shorter clips
Correct Answer: The frame sampling rate is too low; increase sampling density or segment the video into shorter clips
Explanation:
Models sample frames at a fixed rate; fast events between sampled frames are missed. Increasing frame density or processing shorter segments preserves temporal detail. Muting audio, shrinking the context, or downscaling do not address the temporal sampling gap.
Incorrect! Try again.
49A speech application must transcribe overlapping speakers, attribute utterances to each, and detect the emotional tone. Which combination of capabilities is required?
Audio Processing
Hard
A.Text-to-speech synthesis plus sentiment lexicon lookup
B.Noise suppression plus fixed-vocabulary command recognition
C.Speech-to-text with speaker diarization plus paralinguistic/emotion analysis
D.Audio classification of genre plus keyword spotting only
Correct Answer: Speech-to-text with speaker diarization plus paralinguistic/emotion analysis
Explanation:
Attributing utterances to speakers requires diarization, transcription requires speech-to-text, and tone detection requires paralinguistic/emotion analysis. TTS, genre classification, and command recognition do not solve the transcription-plus-attribution-plus-emotion task.
Incorrect! Try again.
50A model is shown a chart image and a paragraph of text that contradict each other about a sales figure. When asked for the correct value, the model should ideally do what to demonstrate robust cross-modal reasoning?
Cross-Modal Reasoning
Hard
A.Average the two conflicting values to produce a compromise figure
B.Always trust the text since language is its primary training modality
C.Surface the discrepancy explicitly and reason about which source is more reliable rather than silently picking one
D.Always trust the image since visual data is inherently more precise
Correct Answer: Surface the discrepancy explicitly and reason about which source is more reliable rather than silently picking one
Explanation:
Robust cross-modal reasoning involves detecting and reporting conflicts between modalities and justifying which to trust. Blindly favoring one modality or averaging numeric conflicts ignores context and can produce a value present in neither source.
Incorrect! Try again.
51A production multimodal assistant must minimize latency and cost while handling both quick text queries and occasional heavy video-analysis requests. Which architectural pattern best balances these goals?
Multimodal Application Design
Hard
A.Route requests to different model tiers/endpoints based on modality and complexity (model routing)
B.Convert all inputs to text first and use only a small text model
C.Cache the first response and return it for all subsequent requests
D.Send every request to the largest multimodal model to maximize accuracy
Correct Answer: Route requests to different model tiers/endpoints based on modality and complexity (model routing)
Explanation:
Model routing sends lightweight queries to cheaper/faster models and reserves large multimodal models for heavy tasks, optimizing cost and latency. Using the largest model for everything wastes resources, text-only conversion loses modality, and blanket caching returns wrong answers.
Incorrect! Try again.
52A developer needs the Gemini model to return strictly parseable output that conforms to a predefined schema for downstream automation. Which approach is the most reliable?
Gemini Models
Hard
A.Ask politely in the prompt to "please return valid JSON" and parse best-effort
B.Set candidateCount high and pick whichever response happens to parse
C.Raise the temperature so the model considers more formatting options
D.Use controlled/structured output with a response schema (e.g., JSON mode with a defined schema)
Correct Answer: Use controlled/structured output with a response schema (e.g., JSON mode with a defined schema)
Explanation:
Constrained decoding against a schema guarantees the output conforms to the required structure. Polite prompting is unreliable, higher temperature increases formatting variance, and generating many candidates hoping one parses is wasteful and non-deterministic.
Incorrect! Try again.
53A team wants to compare a fine-tuned model against the base model on a held-out set and only promote the challenger if it improves a business metric by a statistically meaningful margin. Which workflow construct best encodes this gate?
End-to-End AI Development Workflow
Hard
A.A cron job that redeploys the newest model version every night unconditionally
B.A conditional pipeline step that runs evaluation and branches to registration only if the metric threshold is met
C.A load balancer weight change made directly on the serving endpoint
D.A manual spreadsheet review performed after the model is already live
Correct Answer: A conditional pipeline step that runs evaluation and branches to registration only if the metric threshold is met
Explanation:
A conditional (gated) pipeline step evaluates the challenger and only proceeds to registration/promotion when thresholds pass, encoding the decision automatically. Unconditional redeploys, post-hoc manual reviews, and raw endpoint edits bypass the evaluation gate.
Incorrect! Try again.
54A user increases the guidance scale (CFG) very high to force stronger adherence to a detailed prompt. What is the most likely trade-off they will observe?
Text-to-Image Generation
Hard
A.Complete elimination of prompt-following in favor of random content
B.Reduced image diversity and possible oversaturation or artifacts despite tighter prompt adherence
C.Guaranteed photorealism with no visual artifacts of any kind
D.Faster generation because fewer denoising steps are needed
Correct Answer: Reduced image diversity and possible oversaturation or artifacts despite tighter prompt adherence
Explanation:
Very high guidance scales push the model hard toward the prompt, reducing diversity and often introducing oversaturation or artifacts. It does not guarantee flawless realism, does not reduce step count, and does not abandon prompt-following.
Incorrect! Try again.
55For an OCR-heavy document understanding task where layout matters (tables, columns), which failure is most characteristic of a general multimodal model versus a specialized layout-aware pipeline?
Image Understanding
Hard
A.Producing perfectly ordered output but wrong language
B.Inability to recognize any characters at all
C.Reading-order errors that scramble multi-column or tabular content
D.Refusing to process any image containing text
Correct Answer: Reading-order errors that scramble multi-column or tabular content
Explanation:
General multimodal models can read text but often infer reading order incorrectly for complex layouts, scrambling columns or table cells. They do recognize characters, do not systematically misidentify language, and do not refuse text images.
Incorrect! Try again.
56A developer wants a Gemini model to answer "at what timestamp does the speaker mention pricing?" for an uploaded video. What must the design account for to make timestamp answers reliable?
Video Analysis
Hard
A.Uploading only a single thumbnail to reduce processing overhead
B.Requesting the answer at temperature to encourage precise timing
C.Removing the audio track since only visual frames carry timing information
D.Providing/aligning temporal references so the model can map content to time offsets in the sampled frames/audio
Correct Answer: Providing/aligning temporal references so the model can map content to time offsets in the sampled frames/audio
Explanation:
Timestamp grounding requires the model to associate content with temporal positions from sampled frames/audio. Removing audio discards spoken-pricing cues, high temperature harms precision, and a single thumbnail contains no temporal information.
Incorrect! Try again.
57An audio understanding system must handle a -minute podcast but the model has a token budget that the full audio exceeds. Which strategy preserves the most useful information within the constraint?
Audio Processing
Hard
A.Increase the sample rate so the audio compresses into fewer tokens
B.Convert the audio to a single image spectrogram and discard the waveform
C.Chunk the audio into overlapping segments, process each, then aggregate and summarize across chunks
D.Truncate the audio to the first few minutes that fit in the budget
Correct Answer: Chunk the audio into overlapping segments, process each, then aggregate and summarize across chunks
Explanation:
Overlapping chunking with cross-chunk aggregation covers the whole podcast while respecting token limits, and overlap preserves context at boundaries. Truncation loses most content, raising sample rate increases (not decreases) data, and a lone spectrogram loses speech detail.
Incorrect! Try again.
58In contrastive multimodal embedding models (e.g., CLIP-style), image and text are projected into a shared space. Which statement best characterizes what "alignment" means and a key limitation?
Cross-Modal Reasoning
Hard
A.Semantically matching image-text pairs have high cosine similarity; but fine-grained compositional relations may be poorly captured
B.Alignment guarantees the model can generate new images directly from the embedding without a decoder
C.Images and text share identical raw vectors; so no training is needed to compare them
D.Every pixel is mapped to a unique word; so any image can be losslessly reconstructed from its text embedding
Correct Answer: Semantically matching image-text pairs have high cosine similarity; but fine-grained compositional relations may be poorly captured
Explanation:
Contrastive training pulls matching pairs together in embedding space (high cosine similarity), enabling cross-modal retrieval, yet such models often struggle with compositional/relational nuances (e.g., attribute binding). The other options misstate the mechanism or overclaim capabilities.
Incorrect! Try again.
59A multimodal medical-imaging assistant is being designed. Beyond accuracy, which design decision is most critical for safe and responsible deployment?
Multimodal Application Design
Hard
A.Maximizing model size so outputs are always trusted without review
B.Auto-approving high-confidence outputs to reduce clinician workload
C.Hiding confidence scores to avoid confusing end users
D.Adding human-in-the-loop review and clear scope/uncertainty disclosure rather than presenting outputs as diagnoses
Correct Answer: Adding human-in-the-loop review and clear scope/uncertainty disclosure rather than presenting outputs as diagnoses
Explanation:
In high-stakes domains, human oversight and transparent uncertainty/scope communication are essential for safety. Larger models are not inherently trustworthy, hiding confidence removes decision signal, and auto-approving outputs removes the safeguard clinicians provide.
Incorrect! Try again.
60A regulated enterprise must ensure that data sent to Vertex AI Gemini endpoints stays within a specific region and is not used to train Google's foundation models. Which combination of controls addresses these requirements?
Overview of Vertex AI
Hard
A.Disable IAM entirely so the pipeline can access data faster
B.Enable higher temperature and store all prompts in a public bucket for auditing
C.Configure data residency via regional endpoints and rely on the enterprise data-governance guarantees that prompts are not used for foundation-model training
D.Use only the Generative AI Studio playground since it encrypts nothing by design
Correct Answer: Configure data residency via regional endpoints and rely on the enterprise data-governance guarantees that prompts are not used for foundation-model training
Explanation:
Regional endpoints enforce data residency, and Vertex AI's enterprise governance provides that customer prompts are not used to train the foundation models. The other options describe insecure or irrelevant misconfigurations that violate the stated requirements.
Incorrect! Try again.
Did this save you a night before the exam?
LPU Notes is free, and it stays free. Ads cover part of the server bill.
The rest comes out of a student's own pocket: the domain, the storage,
and keeping the site up through the weeks everyone needs it at once.
The payment button didn't load. An ad blocker or a filtered network is the usual reason.
to try again.
Nothing here is ever locked, and nothing unlocks. Chip in only if it was worth it.
What it pays for →