Showing posts with label Machine learning. Show all posts
Showing posts with label Machine learning. Show all posts

Wednesday, 4 September 2019

Introduction to Machine Learning: Assignment 5

1. Which of the following are false?

a. If a linear separable decision boundary exists for a classification problem, the perceptron model
is capable of finding it
b. One perceptron can be trained with zero training error for an XOR function
c. The backpropagation algorithm updates the parameters using gradient descent rule
d. While training a neural network for binary classification task, an ideal choice for the initialization
of parameters should be large random numbers so that the gradient is higher

Ans: b and d

2. For training a binary classification model with three independent variables, you choose to
use neural networks. You apply one hidden layer with four neurons. What are the number of parameters to be estimated? (Consider the bias term as a parameter)

a. 16
b. 21
c. 3^4 = 81
d. 4^3 = 64
e. 12
f. 4
g. None of these

Ans: e

3. Consider the following function f(x) = exp(x)/(1+exp(x))
The derivative f '(x) will be:

Ans: f(x).ln(1-f(x))

4. Suppose the marks obtained by randomly sampled students follow a normal distribution with
unknown . A random sample of 5 marks are 30, 50, 69, 23 and 99. Using the given samples find the maximum likelihood estimate for the mean

a. 54.2
b. 67.75
c. 50
d. Information not sufficient for estimation

Ans: b

5. Some points are sampled from a Probability distribution following the given probability
distribution function Gaussian with variance x
where x > 0 and m > 0. The collected points are 10, 12, 16, 14 and 15. Give the maximum
likelihood estimate for m.
a. 13.4
b. 13.02
c. 14
d. 20
e. None of these

Ans: b

6. We have a function which takes a two-dimensional input x=(x1,x2) and has two parameters w=(w1,w2) given by f(x)=                      where sigma(x)1 /(1+exp(-x)). We use back propagation to
estimate the right parameter values.
We start by setting both the parameters to 1. Assume that we are given a training point . Given this information answer the next two questions. What is the value of  derivative of f w.r.t w2?

a. 0.098
b. 0.693
c. 0.143
d. -0.367

Ans: b

7. In the previous question, if the learning rate is 0.5, what will be the value of after one
update using backpropagation algorithm?

a. -0.4423
b. 0.4423
c. 1.62
d. 0.381

Ans: b

8. Which of the following is NOT a valid conjugate prior?

a. Gaussian - Gaussian
b. Beta - Binomial
c. Binomial - Bernoulli
d. Beta - Bernoulli

Ans: c

Wednesday, 21 August 2019

Introduction to Machine Learning: Assignment 3

1. Given that the decision boundary separating two classes is linear, what can be inferred about
the discriminant functions of the two classes?

a. Both discriminant functions have to be necessarily linear
b. At least one of the discriminant functions is linear
c. Both discriminant functions can be non-linear

Ans: a

2. We discussed the concept of masking in video lectures. What are the minimum number of
basis transformations required in order to avoid masking for K classes?

a. K
b. K-1
c. K
d. K(K-1)/2

Ans: a

3. Consider the function f1(x) and f2(x) shown in the figure below
Which of the following is correct?

Ans: a: 0 < beta < alpha

4. Which of the following is correct about linear discriminant analysis ?

a. It minimizes the variance between the classes relative to the within class variance
b. It maximizes the within class variance relative to the variance between classes
c. It maximizes the variance between the classes relative to the within class variance
d. None of these

Ans: c

5. Consider the case where two classes follow Gaussian distribution which are centered at
(−1, 2) and (1, 4) and have identity covariance matrix. Which of the following is the separating decision boundary?

(a) y − x = 3
(b) x + y = 3
(c) x + y = 6
(d) (b) and (c) are possible
e) None of these
(f) Can not be found from the given information

Ans: c

6. Consider the following data with two classes. The color indicates different class.
Which of the following models (with NO additional complexity) can achieve zero training error for
classification?

a. LDA
b. PCA
c.  Logistic regression
d. None of these

Ans: c

7. Which of the following technique is well known to utilize class labels in feature selection for
classification?

(a) LDA
(b) PCA (by dimensionality reduction)
(c) both (a) and (b)
(e) None of these

Ans: a

8. We discussed the use of MLE for the estimation of parameters of logistic regression model.
We used which of the following assumptions to derive the likelihood function ?

a. independence among the class labels
b. independence among each training sample
c. independence among the parameters of the model
d. None of these

Ans: b

9. Consider the following distribution of training data:
Which method would you choose for dimensionality reduction?

(a) Linear Discriminant Analysis
(b) Principal Component Analysis
(c) (a) and (b) perform very poorly, so have to choose Quadratic Discriminant Analysis
(d) (a) or (b) are equally good
(e) None of these

Ans: b

Introduction to Machine Learning: Assignment 2

1. The parameters obtained in linear regression

a. can take any value in the real space
b. are strictly integers
c. always lie in the range [0,1]
d. can take only non zero values

Ans: a

2. Consider forward selection, backward selection and best subset selection with respect to the
same data set. Which of the following is true?

a. Best subset selection can be computationally more expensive than forward selection
b. forward selection and backward selection always lead to the same result
c. best subset selection can be computationally less expensive than backward selection
d. best subset selection and forward selection are computationally equally expensive
e. both (b) and (d)

Ans: a

3. Adding interaction terms (such as products of two dimensions) along with original features in
linear regression

a. can reduce training error
b. can increase training error
c. cannot affect training error

Ans: a

4. Consider the following five training examples X = [2 3 4 5 6]  Y = [12.89 17.75 23.31 28.31 32.13] We want to learn a function of the form which is parameterized by (a,b).Using squared error as the loss function, which of the following parameters would you use to model this function

a. (4 3)
b. (5 3)
c. (5 1)
d. (1 5)

Ans: b

5. A study was conducted to understand the effect of number of hours the students spent
studying to their performance in the final exams. You are given the following 8 samples from the study.What is the best linear fit on this dataset?

a. y = -3.39x + 11.62
b. y = 4.59x + 12.58
c. y = 3.39x + 10.58
d. y = 4.69x + 11.62

Ans: b

6. Which of the following shrinkage method is more likely to lead to sparse solution?
a. Lasso regression
b. Ridge regression
c. Lasso and ridge regression both return sparse solutions

Ans: a

7. Consider the design matrix X of dimension N x (p+1) . Which of the following statements
are true?

a. The row space of X is the same as the column space of X^T
b. The row space of is the same as the row space of X^T
c. both (a) and (b)
d. none of the above

Ans: a

8. How does LASSO differ from Ridge Regression?

a. LASSO uses regularization while Ridge Regression uses regularization
b. LASSO uses regularization while Ridge Regression uses regularization
c. The LASSO constraint is a high-dimensional rhomboid while the Ridge Regression constraint is a
high-dimensional ellipsoid
d. Ridge Regression shrinks more coefficients to 0 compared to LASSO
e. The Ridge Regression constraint is a high-dimensional rhomboid while the LASSO constraint is a
high-dimensional ellipsoid
f. Ridge Regression shrinks less coefficients to 0 compared to LASSO

Ans: a,c and f

9. Principal Component Regression (PCR) is an approach to find an orthogonal set of basis vectors which can then be used to reduce the dimension of the input. Which of the following matrices contains the principal component directions as its columns (follow notation from the lecture video)

a. X
b. S
c. Xc
d. V
e. U

Ans: c

10. Let v , v , . . . v denote the Principal Components of some data X, as extracted by
Principal Components Analysis and where v is the First Principal
Component. What can you say about the variance of X in the directions defined by v , v , . . . v?

a. X has the highest variance along v
b. X has the lowest variance along v
c. X has the lowest variance along v
d. X has the highest variance along v
e. Order of variance : v ≥ v ≥ . . . ≥ v
f. Order of variance : v ≥ v ≥ . . . ≥ v

Ans: a,b and e

Tuesday, 13 August 2019

Introduction to Machine Learning: Assignment 1


1)  Which of the following is a supervised learning problem?
a. Predicting the outcome of a cricket match as win or loss based on historical data. 
b. Recommending a movie to an exisiting user on a website like IMdB based on the search history (including other users)
c. Predicting the gender of a person from his/her image. You are given the data of 1 Million images along the gender
d. Given the class labels of old news articles, predicting the class of a new news article from its content. Class of a news article can be such as sports, politics, technology, etc

Ans: a and c2) Which of the following are classification problems?
a. Predicting the temperature (in Celsius) of a room from other environmental features (such as atmospheric pressure, humidity etc)
b. Predicting if a cricket player is a batsman or bowler given his playing records
c. Finding the shorter route between two existing routes between two points.
d. Predicting if a particular route between two points has traffic jam or not based on the travel time of vehicles
Ans: b and d3) Which of the following is a regression task?
a. Predicting the monthly sales of a cloth store in rupees
b. Predicting if a user would like to listen to a newly released song or not based on historical data
c. Predicting the confirmation probability (in fraction) of your train ticket whose current status is waiting list based on historical data
d. Predicting if a patient has diabetes or not based on historical medical records.
Ans: a and c4) Which of the following is an unsupervised task?
a. Grouping images of footwear and caps separately for a given set of images
b. Learning to play chess
c. Predicting if an edible item is sweet or spicy based on the information of the ingredients and their
quantities.
d. all of the above
Ans: a
5) Which of the following is a categorical feature?
a. Number of legs of an animal
b. Number of hours you study in a day
c. Branch of an engineering student
d. Your weekly expenditure in rupees.
Ans: c
6) Let X and Y be a uniformly distributed random variable over the interval [0,4] and [0,3]
respectively. If X and Y are independent events, then compute the probability, P(max(X,Y)>2)

a. 1/6
b. 5/6
c. 2/3
d. None of the above
Ans: b or c [doubt]
7) Let the trace and determinant of a matrix   [a b; c d] be 4 and 3 respectively. The
eigenvalues of A are

a. [3+sqrt(7i)]/2 , [3+sqrt(7i)]/2, where i = sqrt(-1)
b. 1,3
c. None of the above
d. Cannot be computed as the entries of the matrix A are not given
Ans: b
8) What would be the ideal complexity of the curve which can be used for polynomial curve
fitting for the data shown below. (y-axis denotes the dependent variable)

a) Linear
b) Quadratic
c) Cubic
d) If there are N training samples, fit a (N − 1) order polynomial for achieving minimum training error
Ans: c
9) Which of the following are true about bias and variance of overfitted and underfitted models?
a) Underfitted models have low bias
b) Underfitted models have high bias
c) Overfitted models have low variance
d) Overfitted models have high variance
Ans: b and d10) What happens when your model complexity increases?
a) Model Bias increases
b) Model Bias decreases
c) Variance of the model increases
d) Variance of the model decreases
Ans: b and c