Unit 5: Natural Language Processing with Sequence and Attention Models - Practice Quiz

INT344 — Natural Language Processing 50 Questions
0 Correct 0 Wrong 50 Left
0/50

1 What is the primary characteristic of a Sequence Model compared to a standard Feedforward Neural Network?

A. It assumes all inputs are independent of each other
B. It takes the order of inputs into account and can handle variable-length inputs
C. It processes inputs of fixed length only
D. It uses only Convolutional layers

2 Which of the following data types is best suited for a Sequence Model?

A. Tabular housing price data
B. Static image classification
C. Sentiment analysis of movie reviews
D. Iris flower categorization

3 In a Recurrent Neural Network (RNN), what is the function of the 'hidden state'?

A. To reset the network weights after every epoch
B. To store the final output class
C. To visualize the attention weights
D. To act as a memory that captures information about previous time steps

4 What is the phenomenon called when the gradients become extremely small during the backpropagation through time in an RNN, preventing weights from updating?

A. Exploding Gradient
B. Vanishing Gradient
C. Overfitting
D. Gradient Clipping

5 Which algorithm is typically used to train Recurrent Neural Networks?

A. Random Forest
B. Backpropagation Through Time (BPTT)
C. Standard Backpropagation
D. K-Means Clustering

6 Which activation function is most commonly used for the hidden state in a simple RNN to help regulate values?

A. Linear
B. Softmax
C. Tanh
D. ReLU

7 What is the primary architectural solution designed to solve the Vanishing Gradient problem in standard RNNs?

A. Autoencoder
B. Convolutional Neural Network (CNN)
C. Perceptron
D. Long Short-Term Memory (LSTM)

8 In an LSTM unit, which gate is responsible for deciding what information to discard from the cell state?

A. Update Gate
B. Forget Gate
C. Input Gate
D. Output Gate

9 What represents the 'long-term memory' component in an LSTM architecture?

A. Cell State (C_t)
B. Input Gate
C. Hidden State (h_t)
D. Output Gate

10 In an LSTM, what is the range of values output by the sigmoid activation function used in gates?

A. 0 to 100
B. -infinity to +infinity
C. 0 to 1
D. -1 to 1

11 Which task involves assigning a grammatical category (like Noun, Verb, Adjective) to every word in a sentence?

A. Sentiment Analysis
B. Named Entity Recognition
C. Machine Translation
D. Part-of-Speech (POS) Tagging

12 Named Entity Recognition (NER) is primarily concerned with identifying:

A. Real-world objects like people, organizations, and locations
B. Sentiment of the text
C. Grammatical errors
D. The translation of the text

13 What type of Sequence problem is POS Tagging?

A. One-to-Many
B. Many-to-Many (Synced)
C. Many-to-One
D. One-to-One

14 In the context of NER, what does the 'BIO' or 'IOB' tagging scheme stand for?

A. Basic-Input-Operation
B. Binary-Input-Output
C. Beginning-Inside-Outside
D. Backward-Inward-Onward

15 What is the core architecture used in Neural Machine Translation (NMT) before the introduction of Attention?

A. Random Forest
B. Encoder-Only
C. Decoder-Only
D. Encoder-Decoder (Seq2Seq)

16 In a traditional Seq2Seq model, what is the role of the Encoder?

A. To generate the output sequence
B. To visualize the data
C. To compress the input sequence into a fixed-length context vector
D. To calculate the loss function

17 What is the 'Context Vector' in a traditional RNN-based Encoder-Decoder model?

A. The average of all input vectors
B. The weights of the output layer
C. The last hidden state of the encoder
D. The first hidden state of the encoder

18 Which of the following is a major bottleneck of the traditional Seq2Seq model?

A. It requires too much RAM
B. Performance degrades significantly for long sentences due to the fixed-length context vector
C. It can only translate into English
D. It cannot handle numeric data

19 What is 'Teacher Forcing' in the context of training sequence models?

A. Manually setting the weights of the network
B. Using the model's predicted output as input for the next step during training
C. Forcing the model to stop training early
D. Using the actual ground truth output from the previous time step as input for the current step during training

20 Which search strategy explores multiple possible output sequences simultaneously to find the most likely translation?

A. Beam Search
B. Linear Search
C. Greedy Search
D. Binary Search

21 The Attention Mechanism was primarily introduced to solve which problem?

A. The inability of RNNs to process images
B. Overfitting in CNNs
C. The information bottleneck of the fixed-length context vector in NMT
D. Slow training of Linear Regression

22 How does the Attention Mechanism calculate the context vector for each time step in the decoder?

A. By averaging all input words
B. By computing a weighted sum of all encoder hidden states
C. By taking the last state of the encoder only
D. By randomly selecting an encoder state

23 In Attention, what do the 'alignment scores' (or attention weights) represent?

A. The number of hidden layers
B. The magnitude of the gradient
C. The error rate of the model
D. How relevant a specific input word is to the word currently being generated

24 What mathematical function is typically applied to alignment scores to convert them into probabilities that sum to 1?

A. Softmax
B. Sigmoid
C. ReLU
D. Tanh

25 In the context of RNNs, what is 'Backpropagation Through Time' (BPTT)?

A. Unfolding the RNN across time steps and applying backpropagation
B. Using future data to predict past data
C. Training the network in reverse order
D. A method to predict future stock prices

26 Which of the following is NOT a gate in a standard LSTM?

A. Attention Gate
B. Input Gate
C. Forget Gate
D. Output Gate

27 What is the shape of the input data for a basic RNN layer in Keras/TensorFlow?

A. (Features, Labels)
B. (Batch Size, Timesteps, Features)
C. (Timesteps, Batch Size)
D. (Batch Size, Features)

28 Why are Bidirectional RNNs (BiRNNs) useful?

A. They use fewer parameters
B. They allow the network to have context from both the past and the future
C. They train faster than standard RNNs
D. They eliminate the need for backpropagation

29 In a Many-to-One sequence model (e.g., Sentiment Analysis), where is the output typically taken?

A. Randomly sampled
B. At every time step
C. At the last time step
D. At the first time step

30 What does the 'candidate cell state' in an LSTM do?

A. It outputs the final prediction
B. It decides what to forget
C. It clears the memory
D. It proposes new values that could be added to the state

31 Which issue leads to the 'Exploding Gradient' problem?

A. Gradients < 1 accumulating multiplicatively
B. Learning rate being too low
C. Gradients > 1 accumulating multiplicatively
D. Weights initialized to zero

32 A solution to the Exploding Gradient problem is:

A. Gradient Clipping
B. Removing the hidden layer
C. Using ReLU
D. Increasing the learning rate

33 In an Attention model, the vector c_t is often referred to as:

A. The Context Vector
B. The Bias Vector
C. The Forget Vector
D. The Noise Vector

34 Sequence-to-Sequence models are most commonly associated with:

A. Cluster Analysis
B. Text Summarization
C. Linear Regression
D. Image Segmentation

35 What is 'Global Attention'?

A. Attention that considers all hidden states of the encoder
B. Attention applied without weights
C. Attention applied to a single word
D. Attention that considers only a window of hidden states

36 Which of the following describes 'Greedy Decoding'?

A. Choosing the word with the highest probability at each step immediately
B. Waiting until the end to choose words
C. Choosing a random word based on distribution
D. Considering all possible future sequences

37 In POS tagging, if a word is ambiguous (e.g., 'book' can be a noun or verb), how does an RNN resolve it?

A. It always picks the most common usage
B. It uses the context provided by surrounding words stored in the hidden state
C. It cannot resolve ambiguity
D. It flips a coin

38 What is the typical loss function for a multi-class classification problem like POS Tagging or NMT?

A. Hinge Loss
B. Mean Squared Error (MSE)
C. Absolute Error
D. Categorical Cross-Entropy

39 What does GRU stand for?

A. Gradient Rectified Unit
B. Gated Recurrent Unit
C. General Regression Unit
D. Global Recurrent Update

40 The 'Input Gate' in an LSTM is usually controlled by which activation function?

A. Sigmoid
B. Linear
C. Tanh
D. ReLU

41 What visual tool is often used to interpret what an Attention model has learned?

A. Attention Heatmap
B. Scatter Plot
C. Histogram
D. Pie Chart

42 Which limitation of RNNs prevents parallelization during training?

A. Sequential dependency of the hidden state
B. Complex loss functions
C. Use of sigmoid functions
D. Large memory footprint

43 In a sequence model, 'padding' is used to:

A. Remove stopwords
B. Add noise to the data
C. Make all sequences in a batch the same length
D. Increase the learning rate

44 Which of these is a 'many-to-one' application of sequence models?

A. Music Generation
B. Sentiment Classification
C. Machine Translation
D. Video Captioning

45 In an NER task, identifying 'Apple' as an Organization rather than a Fruit relies on:

A. The contextual information in the sequence
B. The capitalization
C. The spelling of the word
D. The length of the word

46 Why is the traditional Encoder-Decoder model often described as having 'amnesia'?

A. It struggles to retain information from the beginning of a long sequence at the decoding stage
B. It uses a forget gate
C. It forgets the weights after training
D. It cannot learn new words

47 In the attention equation score(h_t, h_s), what are h_t and h_s?

A. Weight and Bias
B. Learning rate and Loss
C. Input and Output Gates
D. Decoder hidden state and Encoder hidden state

48 Which mechanism allows a model to focus on 'local' parts of the input sequence based on the current decoding step?

A. Dropout
B. Attention Mechanism
C. Batch Normalization
D. Max Pooling

49 In sequence labeling, what does the output layer usually consist of?

A. A linear regression layer
B. A Softmax layer over the tag set for each time step
C. A single neuron
D. A clustering algorithm

50 Which of the following best describes the 'Seq2Seq' mapping?

A. Variable Input Size -> Fixed Output Size
B. Fixed Input Size -> Variable Output Size
C. Fixed Input Size -> Fixed Output Size
D. Variable Input Size -> Variable Output Size