Unit 6: Introduction to GPU Architecture - Practice Quiz

CSE211 — Computer Organization And Design 60 Questions
0 Correct 0 Wrong 60 Left
0/60

1 What does the acronym GPU stand for?

Nvidia Case Study Easy
A. Graphical Program Unit
B. General Processing Unit
C. Global Processing Utility
D. Graphics Processing Unit

2 Which parallel computing platform and programming model was developed by Nvidia?

Nvidia Case Study Easy
A. OpenCL
B. DirectX
C. CUDA
D. Vulkan

3 In Nvidia GPU terminology, the basic parallel execution units are called:

Nvidia Case Study Easy
A. Register banks
B. CUDA cores
C. Logic gates
D. Cache lines

4 GPUs are especially well suited for which type of processing?

Nvidia Case Study Easy
A. Sequential processing
B. Single-thread processing
C. Interrupt handling
D. Parallel processing

5 A supercomputer is a computer that primarily offers very high:

Introduction to Supercomputer Easy
A. Computational performance
B. Storage cost
C. Battery life
D. Display resolution

6 The performance of supercomputers is commonly measured in:

Introduction to Supercomputer Easy
A. FLOPS
B. Pixels
C. Bytes
D. Hertz only

7 Which term describes a supercomputer capable of at least floating-point operations per second?

Introduction to Supercomputer Easy
A. Terascale
B. Exascale
C. Petascale
D. Gigascale

8 Most modern supercomputers achieve their power by using:

Introduction to Supercomputer Easy
A. Analog circuits
B. Massively parallel architectures
C. A single fast CPU
D. Mechanical relays

9 What is the basic unit of information in quantum computing called?

Introduction to Qubits and Quantum Computing Easy
A. Nibble
B. Byte
C. Bit
D. Qubit

10 Unlike a classical bit, a qubit can exist in a combination of both states at once due to:

Introduction to Qubits and Quantum Computing Easy
A. Compression
B. Pipelining
C. Multiplexing
D. Superposition

11 The quantum phenomenon where two qubits become correlated so the state of one affects the other is called:

Introduction to Qubits and Quantum Computing Easy
A. Refraction
B. Entanglement
C. Modulation
D. Diffusion

12 A classical bit can hold how many possible values at any given time?

Introduction to Qubits and Quantum Computing Easy
A. One only
B. Infinite values
C. Ten values
D. Two ( or )

13 Which trend refers to placing multiple processing cores on a single chip?

Latest Technology and Trends in Computer Architecture Easy
A. Virtualization
B. Emulation
C. Multicore processing
D. Overclocking

14 Moore's Law originally predicted that the number of transistors on a chip would:

Latest Technology and Trends in Computer Architecture Easy
A. Halve every year
B. Increase only once per decade
C. Stay constant over time
D. Roughly double every couple of years

15 Which term describes processors that combine CPU and GPU on the same die?

Latest Technology and Trends in Computer Architecture Easy
A. Analog computing
B. Heterogeneous computing
C. Serial computing
D. Homogeneous storage

16 Which open-standard instruction set architecture is widely used in next-generation designs?

Next Generation Processors Architecture Easy
A. x86-legacy
B. COBOL
C. HTML
D. RISC-V

17 Executing more than one instruction per clock cycle is a feature of which processor type?

Next Generation Processors Architecture Easy
A. Superscalar
B. Single-cycle
C. Sequential-only
D. Serial

18 The term microarchitecture refers to:

Microarchitecture Easy
A. The operating system kernel
B. How a processor implements its instruction set
C. The physical case of a computer
D. The user interface design

19 Which microarchitectural technique overlaps the execution stages of multiple instructions?

Microarchitecture Easy
A. Defragmentation
B. Formatting
C. Pipelining
D. Encryption

20 Which company designs the ARM-based A-series and M-series chips used in its smartphones, tablets, and desktops?

Latest Processor for Smartphone or Tablet and Desktop Easy
A. Qualcomm
B. Apple
C. Intel
D. AMD

21 In NVIDIA's CUDA programming model, a group of threads that execute the same instruction in lockstep on an SM is called a:

Nvidia Case Study Medium
A. Grid
B. Kernel
C. Block
D. Warp

22 NVIDIA introduced dedicated units for accelerating matrix multiply-accumulate operations used in deep learning. These are known as:

Nvidia Case Study Medium
A. Texture Units
B. Vertex Shaders
C. Raster Units
D. Tensor Cores

23 If an NVIDIA GPU has 80 Streaming Multiprocessors and each SM can hold up to 2048 resident threads, the maximum number of concurrently resident threads is:

Nvidia Case Study Medium
A. 163840
B. 81920
C. 2048
D. 160000

24 The performance of supercomputers in the TOP500 list is primarily measured using which benchmark?

Introduction to Supercomputer Medium
A. SPECint (IPC)
B. Geekbench (score)
C. Dhrystone (DMIPS)
D. LINPACK (FLOPS)

25 A supercomputer achieves floating-point operations per second. This performance level is classified as:

Introduction to Supercomputer Medium
A. Petascale
B. Gigascale
C. Exascale
D. Terascale

26 Modern supercomputers largely achieve high performance through which architectural approach?

Introduction to Supercomputer Medium
A. Larger cache on a single monolithic chip
B. Increasing clock speed of one CPU core
C. Single very high frequency serial processor
D. Massively parallel processing with many interconnected nodes

27 A quantum system with qubits can represent how many basis states simultaneously in superposition?

Introduction to Qubits and Quantum Computing Medium
A.
B.
C.
D.

28 Which property allows two qubits to be correlated such that the state of one instantly influences the measured state of the other?

Introduction to Qubits and Quantum Computing Medium
A. Tunneling
B. Decoherence
C. Superposition
D. Entanglement

29 The loss of quantum information due to unwanted interaction with the environment is called:

Introduction to Qubits and Quantum Computing Medium
A. Measurement
B. Superposition
C. Decoherence
D. Interference

30 A single qubit state must satisfy which normalization condition?

Introduction to Qubits and Quantum Computing Medium
A.
B.
C.
D.

31 The trend of building processors from multiple smaller dies connected within a single package is known as:

Latest Technology and Trends in Computer Architecture Medium
A. Frequency boosting
B. Deep pipelining
C. Monolithic scaling
D. Chiplet design

32 The end of Dennard scaling primarily led to which shift in processor design?

Latest Technology and Trends in Computer Architecture Medium
A. Moving to multi-core parallelism instead of raising clock frequency
B. Abandoning cache hierarchies entirely
C. Reducing the number of cores per chip
D. Returning to single-core CISC designs

33 High Bandwidth Memory (HBM) achieves greater bandwidth than traditional DDR mainly by:

Latest Technology and Trends in Computer Architecture Medium
A. Placing memory on a separate PCB far from the CPU
B. Increasing the DRAM clock frequency alone
C. Stacking DRAM dies vertically with a wide interface
D. Using SRAM cells instead of DRAM

34 Domain-specific accelerators (like TPUs) outperform general-purpose CPUs for AI workloads mainly because they:

Next Generation Processors Architecture Medium
A. Use only single-threaded execution
B. Have larger general-purpose instruction sets
C. Optimize hardware for specific operations like matrix math
D. Run at much higher clock frequencies

35 The RISC-V instruction set architecture is notable for being:

Next Generation Processors Architecture Medium
A. A closed CISC design for mainframes
B. A proprietary ISA owned by Intel
C. A GPU-only shader language
D. An open and royalty-free ISA

36 Neuromorphic processors are designed to mimic which system to achieve energy-efficient computation?

Next Generation Processors Architecture Medium
A. The bus topology of a network switch
B. The layout of a mechanical hard drive
C. The biological neural structure of the brain
D. The pipeline of a classic RISC CPU

37 In a superscalar processor, out-of-order execution improves performance primarily by:

Microarchitecture Medium
A. Executing independent instructions as their operands become ready
B. Lowering the processor clock frequency
C. Reducing the total number of instructions in a program
D. Removing the need for a register file

38 A pipeline with 5 stages ideally increases throughput compared to a non-pipelined design by a factor of up to:

Microarchitecture Medium
A. 5
B. 10
C. 25
D. 1

39 Branch prediction is used in modern microarchitectures mainly to:

Microarchitecture Medium
A. Reduce stalls caused by control hazards
B. Lower the number of pipeline stages
C. Increase the size of the register file
D. Eliminate cache misses

40 Simultaneous Multithreading (SMT), such as Intel's Hyper-Threading, improves utilization by:

Microarchitecture Medium
A. Running each thread on a separate clock domain
B. Increasing the L1 cache size automatically
C. Issuing instructions from multiple threads to shared execution units
D. Doubling the physical number of CPU cores

41 In NVIDIA's SIMT execution model, a warp of 32 threads encounters a data-dependent branch where 20 threads take the if path and 12 take the else path. Assuming no independent thread scheduling (pre-Volta), what is the effective SIMD utilization during the divergent region?

Nvidia Case Study Hard
A. The warp splits into two independent warps of 20 and 12 threads with full utilization each
B. Only the majority path executes, discarding the 12 minority threads permanently
C. The warp executes both paths serially, so utilization drops to roughly averaged across the two masked passes
D. Both paths execute in parallel using two separate ALU lanes, keeping utilization at

42 NVIDIA Tensor Cores accelerate matrix multiply-accumulate operations. If a Tensor Core computes a by matrix multiply-accumulate () per clock, how many fused multiply-add (FMA) operations does one such operation represent?

Nvidia Case Study Hard
A. FMAs (i.e., )
B. FMAs (i.e., )
C. FMAs (i.e., )
D. FMAs (i.e., )

43 A supercomputer achieves a peak (theoretical) performance of PFLOPS but a sustained LINPACK () of PFLOPS. What does the ratio primarily quantify, and what is its approximate value here?

Introduction to Supercomputer Hard
A. Interconnect bandwidth utilization; about
B. Power efficiency in FLOPS/watt; about
C. Computational efficiency; about
D. Memory bound fraction; about

44 Amdahl's Law limits the speedup of a parallel program. If of a workload is parallelizable and a supercomputer offers effectively unlimited processors, what is the maximum theoretical speedup?

Introduction to Supercomputer Hard
A.
B. Unbounded (infinite)
C.
D.

45 A quantum register of qubits can represent a superposition of how many basis states simultaneously, and how does this scale compared to classical bits?

Introduction to Qubits and Quantum Computing Hard
A. states simultaneously, versus states for classical bits
B. states simultaneously, versus states for classical bits
C. states simultaneously, versus one of states for classical bits
D. states simultaneously, versus states for classical bits

46 Which statement best characterizes why quantum decoherence is a fundamental engineering challenge, rather than merely a noise problem solvable by shielding?

Introduction to Qubits and Quantum Computing Hard
A. Decoherence increases the number of usable qubits, requiring fewer error-correction resources
B. Decoherence only affects readout accuracy and disappears entirely at absolute zero temperature
C. Interaction with the environment collapses fragile superposition/entanglement, and it grows worse as qubit count and gate depth increase
D. Decoherence is caused solely by cosmic rays and is eliminated by underground placement

47 Grover's search algorithm provides a quadratic speedup for unstructured search over items. For a database of entries, approximately how many oracle queries does Grover's algorithm need versus a classical average-case search?

Introduction to Qubits and Quantum Computing Hard
A. About (order ) versus classically
B. About (order ) versus about classically
C. About versus (no advantage)
D. About versus classically

48 As Dennard scaling ended, power density stopped decreasing with transistor size, leading to the 'dark silicon' phenomenon. What does dark silicon most directly refer to?

Latest Technology and Trends in Computer Architecture Hard
A. The fraction of a chip that must be powered off or run at low frequency at any time due to thermal/power limits
B. Cache regions permanently disabled because of manufacturing defects
C. Transistors fabricated with light-blocking coatings to reduce photonic interference
D. The portion of the die dedicated to analog rather than digital circuits

49 Chiplet-based designs (e.g., disaggregating a monolithic die into multiple smaller dies on an interposer) primarily improve manufacturing economics because of which relationship?

Latest Technology and Trends in Computer Architecture Hard
A. Smaller dies always run at higher clock frequencies, increasing performance per wafer
B. Yield rises as die area shrinks, since defect probability scales with area, so smaller chiplets have far higher per-die yield
C. Larger monolithic dies have higher yield because defects are averaged over more area
D. Chiplets eliminate the need for any inter-die interconnect, reducing latency to zero

50 High Bandwidth Memory (HBM) achieves very high bandwidth despite modest clock rates primarily through which architectural approach?

Latest Technology and Trends in Computer Architecture Hard
A. Eliminating the memory controller and connecting DRAM directly to registers
B. Wide interfaces from vertically stacked DRAM dies connected by through-silicon vias (TSVs), giving thousands of I/O bits per stack
C. Storing data optically rather than electrically to bypass RC delays
D. Extremely high per-pin clock rates exceeding GHz on a narrow 64-bit bus

51 In a heterogeneous big.LITTLE / performance-efficiency core design, the scheduler migrates a latency-sensitive foreground thread from an efficiency core to a performance core. What is the primary architectural challenge that makes such migration non-trivial?

Next Generation Processors Architecture Hard
A. Efficiency cores lack any cache, so all data must be reloaded from disk
B. Maintaining cache and architectural state consistency (ISA compatibility, coherent caches) so the thread resumes correctly across cores
C. The two core types cannot share the same physical memory address space
D. Performance cores use a completely different instruction set requiring binary recompilation

52 RISC-V's use of modular, optional ISA extensions (e.g., M, A, F, D, V) contrasts with fixed ISAs. From an architectural standpoint, what is the main trade-off this modularity introduces?

Next Generation Processors Architecture Hard
A. Elimination of all compiler complexity versus reduced instruction count
B. Guaranteed higher clock speeds versus mandatory larger die area
C. Flexibility and minimal base area versus potential software fragmentation across incompatible extension sets
D. Fixed power consumption versus variable memory bandwidth

53 A superscalar out-of-order processor uses register renaming. What fundamental hazard does register renaming eliminate, and what hazard does it NOT eliminate?

Microarchitecture Hard
A. Eliminates all three of RAW, WAR, and WAW simultaneously
B. Eliminates structural hazards; does not eliminate WAR dependencies
C. Eliminates RAW dependencies; does not eliminate control hazards
D. Eliminates WAR and WAW (false/name dependencies); does not eliminate RAW (true data dependencies)

54 Consider a 5-stage pipeline with a branch resolved in the execute (3rd) stage and no branch prediction (predict-not-taken with squash on taken). If of instructions are branches and of those are taken, causing a 2-cycle penalty each, what is the approximate CPI contribution from branch stalls?

Microarchitecture Hard
A. CPI
B. CPI
C. CPI
D. CPI

55 Simultaneous Multithreading (SMT, e.g., Hyper-Threading) improves throughput by filling idle issue slots from multiple threads. Under what condition does SMT provide the LEAST benefit?

Microarchitecture Hard
A. When a single thread already achieves high issue-slot utilization (few idle slots to fill)
B. When the pipeline is very wide with many unused functional units
C. When threads have many long-latency cache misses creating idle slots
D. When the workload has abundant thread-level parallelism

56 Apple's M-series and A-series SoCs employ a Unified Memory Architecture (UMA). Compared to a traditional discrete-GPU system with separate VRAM, what is the principal architectural advantage of UMA for GPU-CPU workloads?

Latest Processor for Smartphone or Tablet and Desktop Hard
A. It removes the need for any memory controller by using registers as main memory
B. It allows the GPU to run a different instruction set than the CPU for higher clocks
C. It doubles total memory capacity by mirroring data between CPU and GPU pools
D. CPU and GPU share one physical memory pool, eliminating explicit copies across a PCIe bus and reducing latency/energy

57 Modern mobile SoCs integrate a dedicated NPU (Neural Processing Unit) alongside CPU and GPU. Why is a specialized NPU more energy-efficient than running the same neural inference on the CPU or GPU?

Latest Processor for Smartphone or Tablet and Desktop Hard
A. It executes the same general-purpose ISA but with more cache
B. It uses fixed-function MAC arrays and reduced-precision arithmetic tailored to tensor ops, minimizing instruction-fetch and data-movement overhead
C. It runs at a much higher clock frequency than the CPU, finishing sooner
D. It stores the entire neural network in registers, avoiding all memory access

58 On an NVIDIA GPU, memory coalescing combines the global-memory accesses of a warp into fewer transactions. If all 32 threads of a warp access consecutive 4-byte words within a single aligned 128-byte segment, how many memory transactions are ideally required?

Nvidia Case Study Hard
A. One 128-byte transaction
B. Four 32-byte transactions
C. Sixteen 8-byte transactions
D. Thirty-two separate 4-byte transactions

59 A supercomputer is rated at PFLOPS and consumes MW of power. What is its energy efficiency in GFLOPS/watt, and why is this metric increasingly the binding constraint for exascale systems?

Introduction to Supercomputer Hard
A. GFLOPS/W; because memory capacity is the sole exascale limiter
B. GFLOPS/W; because total power and cooling, not transistor count, cap achievable scale
C. GFLOPS/W; because software parallelism is the only constraint
D. GFLOPS/W; because interconnect latency dominates exascale design

60 Near-memory and in-memory computing (e.g., processing-in-memory, PIM) architectures aim to overcome which specific bottleneck of the classical von Neumann model?

Next Generation Processors Architecture Hard
A. The instruction-level parallelism limit imposed by branch mispredictions
B. The memory wall: the growing gap and energy cost of moving data between CPU and memory relative to compute
C. The lack of floating-point units in the memory controller
D. The power wall caused solely by clock frequency limits inside the ALU