This page is meant to give you quick access to problems and their solutions. Refer to the original exam PDF, linked above, for test-taking instructions and formatting. Note that we’ve kept the problem text identical, which is why you may see things like “write your answer in the box below” despite there not being a box on this page.
Consider two datasets, \(A=\lbrace y_1,y_2,\ldots,y_n\rbrace \) and \(B=\lbrace z_1,z_2,\ldots,z_m\rbrace \). Let \(R_A(w)\) and \(R_B(w)\) represent the mean squared errors of the constant prediction \(w\) on datasets \(A\) and \(B\), respectively:
Setting this equal to zero gives \(\boxed{w^{\ast}=10}\). The derivative is negative below \(10\) and positive above \(10\), so this is the minimizer.
This is the same kind of optimization we used to minimize mean squared error: the optimal prediction balances the two squared deviations. The constants \(+5\) and \(+17\) disappear when we take the derivative, so they do not affect the minimizer.
c)
6 pts Let \(M(w)\) return the larger of the two mean squared error functions:
$$ M(w)=\max(R_A(w),R_B(w)) $$
Find \(w^{\ast}\), the value of \(w\) that minimizes \(M(w)\). Give your answer as a number with no variables. Hint: Draw a picture — a clearly annotated picture, accompanied by some algebra, is sufficient justification.
Since \(R_A(w)-R_B(w)=12(w-11)\), when \(w>11\), we have \(R_A(w)>R_B(w)\), so \(M(w)\) returns \(R_A(w)\). Otherwise, \(M(w)\) returns \(R_B(w)\); at \(w=11\), the two are equal. Thus,
\(R_B\) is decreasing for \(w<11\), since its vertex is at \(13\). \(R_A\) is increasing for \(w>11\), since its vertex is at \(7\). Therefore, \(M\) decreases up to \(11\) and increases after \(11\), giving \(\boxed{w^{\ast}=11}\) and \(M(11)=21\).
Problem 2 14 pts
Consider a dataset of seven values \(y_1,y_2,\ldots,y_7\), listed in non-decreasing order:
$$ y_1,\ y_2,\ 47,\ 60,\ 60,\ 81,\ y_7 $$
Suppose that the mean of \(y_1,\ldots,y_6\) — that is, not including \(\mathbf{y_7}\) — is \(50\). Furthermore, suppose
\(\displaystyle f(w)=\frac17\sum_{i=1}^7|y_i-w|\), and \(\alpha^{\ast}\) is the value of \(w\) that minimizes \(f(w)\).
\(\displaystyle g(w)=\frac17\sum_{i=1}^7(y_i-w)^2\), and \(\beta^{\ast}\) is the value of \(w\) that minimizes \(g(w)\).
a)
2 pts What is the value of \(\alpha^{\ast}\)? Give your answer as a number with no variables.
\(\alpha^{\ast}=\_\_\_\_\_\_\)
Solution
Mean absolute error is minimized by the median. Since there are seven ordered values, the median is the fourth value, so \(\boxed{\alpha^{\ast}=60}\). The fact that \(60\) is repeated does not change this result: the middle (fourth) value is still \(60\), so the median is still \(60\).
b)
5 pts What is the largest possible value of \(y_7\) such that \(\alpha^{\ast}\geq\beta^{\ast}\)? Give your answer as a number with no variables.
Solution
Mean squared error is minimized by the mean. The first six values sum to \(6\cdot50=300\), so
$$ \beta^*=\frac{300+y_7}{7} $$
Using \(\alpha^{\ast}=60\), the required condition becomes
The value \(120\) is consistent with the ordering, so the largest possible value is \(\boxed{120}\).
c)
7 pts Consider the same seven values, listed in non-decreasing order:
$$ y_1,\ y_2,\ 47,\ 60,\ 60,\ 81,\ y_7 $$
Recall,
$$ f(w)=\frac17\sum_{i=1}^7|y_i-w| $$
For this part, do not assume that \(\alpha^{\ast}\geq\beta^{\ast}\).
Suppose that \(f(59)=30\). What is \(f(65)\)? Give your answer as a number with no variables. Hint: Consider the slopes of \(f(w)\) between \(w=59\) and \(w=65\).
Solution
For any value of \(w\) that is not equal to a \(y_i\), the slope of mean absolute error, \(f(w)\), is
$$ \frac{\text{d}}{\text{d}w}f(w) =\frac{\text{# of points left of }w-\text{# of points right of }w}{7} $$
For \(59<w<60\), there are three values to the left and four to the right, so the slope is \(-1/7\). For \(60<w<65\), there are five to the left and two to the right, so the slope is \(3/7\). Both copies of \(60\) count.
Consider a dataset of \(n\) points, \((x_1,y_1),\ldots,(x_n,y_n)\), where not all of the \(x_i\) are the same. We’d like to fit a modified simple linear regression model whose slope and intercept are forced to be the same:
$$ h(x_i)=w+wx_i $$
The value of \(w\) that minimizes mean squared error for this model is
\(h^{\ast}(x_i)=w^{\ast}+w^{\ast}x_i\) is the fitted modified model, where \(w^{\ast}\) minimizes mean squared error.
\(g^{\ast}(x_i)=w_0^{\ast}+w_1^{\ast}x_i\) is the regular simple linear regression model fitted to the same dataset, where \(w_0^{\ast}\) and \(w_1^{\ast}\) are chosen to minimize mean squared error.
Fill in the \(\boxed{???}\) with the relationship that is guaranteed:
Think of regular simple linear regression as being more flexible: it can have the same slope and intercept if that minimizes MSE, and it can have different slope and intercept if that minimizes MSE even further. Anything the modified model can do, the regular model can do, and more. So the regular model’s minimum MSE cannot be larger.
b)
6 pts Let \(r\) be the correlation coefficient between the \(x\)- and \(y\)-values, and let \(\sigma_x\) and \(\sigma_y\) be their standard deviations, respectively. Suppose
Give each answer as a number with no variables. If it is not possible to determine an answer from the information given, write N/A in the box.
(3 pts) What is \(w^{\ast}\), the optimal slope for the modified simple linear regression model?
\(w^{\ast}=\_\_\_\_\_\_\)
(3 pts) What is \(w_1^{\ast}\), the optimal slope for the regular simple linear regression model?
\(w_1^{\ast}=\_\_\_\_\_\_\)
Solution
(i)\(\boxed{\text{N/A}}\). Moving the data changes which modified line fits best, even if its shape, correlation, and standard deviations stay the same. Every modified line \(y=w+wx\) passes through \((-1,0)\), so it cannot simply move along with the data.
For example, use \(x\)-values \(-1,-1,7,7\) and \(y\)-values \(-2,14,-14,2\) in the left panel. In the right panel, add \(8\) to every \(y\)-value. Both datasets have \(r=-3/5\), \(\sigma_x=4\), and \(\sigma_y=10\).
At \(x=-1\), the prediction is always \(0\), regardless of \(w\). At \(x=7\), the prediction is \(8w\), so the best choice makes \(8w\) equal to the mean of the two \(y\)-values there. This gives \(w^{\ast}=-3/4\) on the left and \(w^{\ast}=1/4\) on the right.
7 pts Show that \(w^{\ast}\), the value of \(w\) that minimizes mean squared error for the modified simple linear regression model \(h(x_i)=w+wx_i\), is
The objective is a quadratic in \(w\) with a positive coefficient on \(w^2\), so this critical point is its minimum. A second derivative test is not necessary.
Problem 4 11 pts
Suppose \(a\), \(b\), and \(c\) are real numbers, and let
6 pts Now suppose that \(a+b+c=0\) and \(\vec u\neq\vec0\). Find \(\theta\), the angle between \(\vec u\) and \(\vec v\). Give your answer as the cosine inverse of a number (e.g. \(\cos^{-1}(3/4)\)). Hint: \((a+b+c)^2=a^2+b^2+c^2+2ab+2bc+2ca\).
5 pts Find \(\vec p\), the projection of \(\vec x\) onto \(\vec v_1\). Give your answer as a vector with two components and no variables.
Solution
Since \(\vec v_1=6\begin{bmatrix}3\\-4\end{bmatrix}\), projecting onto \(\vec v_1\) is equivalent to projecting onto \(\begin{bmatrix}3\\-4\end{bmatrix}\). So,
Equivalently, using \(\vec v_1\) directly gives the coefficient \(300/900=1/3\).
b)
7 pts Suppose that the first component of \(\vec v_2\) is positive. Write \(\vec x\) as a linear combination of \(\vec v_1\) and \(\vec v_2\) by filling in the two boxes below.
A vector perpendicular to \(\begin{bmatrix}3\\-4\end{bmatrix}\) with positive first component has direction \(\begin{bmatrix}4\\3\end{bmatrix}\). This direction vector has length \(5\), so
5 pts Find an equation for the plane \(P\) of the form \(Ax+By+Cz+d=0\). Simplify your answer so that \(A=1\).
Solution
Because \(P\) is a span, it contains the origin, so \(d=0\). With \(A=1\), the equation is \(x+By+Cz=0\). Substituting the two spanning vectors gives
$$ 12-2B+6C=0,\qquad 3+3C=0 $$
The second equation gives \(C=-1\), and then the first gives \(B=3\). Thus,
$$ \boxed{x+3y-z=0} $$
Both vectors lie in this plane and are linearly independent, so they span the entire plane.
For part b), suppose that \(\vec x_3,\vec x_4\in\mathbb R^3\) and that \(\operatorname{span}(\lbrace \vec x_1,\vec x_2,\vec x_3,\vec x_4\rbrace )\) is a plane in \(\mathbb R^3\).
b)
5 pts For each statement below, determine whether it is impossible, possible (but not guaranteed), or guaranteed to be true, given the above assumptions. The first statement has been done for you as an example.
statement
impossible?
possible?
guaranteed?
\(i\)
\(\lVert \vec x_3 \rVert = 5\).
\(ii\)
\(\vec x_3\) and \(\vec x_4\) are scalar multiples of one another.
\(iii\)
\(\vec x_4 = \vec 0\).
\(iv\)
The vectors \(\vec x_1, \vec x_2, \vec x_3\) are linearly independent.
\(\lbrace \vec x_1, \vec x_2\rbrace \) is a basis for \(P\).
Solution
The span of all four vectors contains \(P\) and is a plane, so it must equal \(P\). Thus, \(\vec x_3\) and \(\vec x_4\) both lie in \(P\).
Possible. A vector in \(P\) can have length \(5\), but it doesn’t need to.
Possible. For example, choose \(\vec x_3=\vec x_1\) and \(\vec x_4=2\vec x_1\). They don’t need to be scalar multiples: we could instead choose \(\vec x_3=\vec x_1\) and \(\vec x_4=\vec x_2\).
Possible. We may choose \(\vec x_4=\vec0\) because the first two vectors already span \(P\).
Impossible. Three vectors in a two-dimensional space cannot be linearly independent.
Possible. Choosing \(\vec x_3=\vec x_1\) and \(\vec x_4=\vec x_2\) works. Choosing both to be zero shows that this is not guaranteed.
Guaranteed. By definition, \(\vec x_1\) and \(\vec x_2\) span \(P\), and they are not scalar multiples of one another, so they form a basis.
Additionally, suppose that \(\vec x_3,\vec x_4\in\mathbb R^3\) and that \(\operatorname{span}(\lbrace \vec x_1,\vec x_2,\vec x_3,\vec x_4\rbrace )\) is a plane in \(\mathbb R^3\).
c)
4 pts Give one nonzero vector orthogonal to \(3\vec x_1-7\vec x_2+12\vec x_3-\vec x_4\). Your answer should have no variables. Briefly justify it.
Solution
From part a), \(\vec n=\begin{bmatrix}1\\3\\-1\end{bmatrix}\) is normal to \(P\). All four vectors lie in \(P\), so \(\vec n\cdot\vec x_i=0\) for \(i=1,2,3,4\). By distributivity,
Thus, one valid answer is \(\boxed{\begin{bmatrix}1\\3\\-1\end{bmatrix}}\). Any nonzero scalar multiple of this vector also works.
For part d), suppose that \(\vec x_3,\vec x_4\in\mathbb R^3\) but that \(\operatorname{span}(\lbrace \vec x_1,\vec x_2,\vec x_3,\vec x_4\rbrace ) = \mathbb R^3\).
d)
5 pts Which statements are impossible? Select all that apply.
\(\vec x_3\) and \(\vec x_4\) are on \(P\).
The vectors \(\vec x_3,\vec x_4\) are linearly independent.
The vectors \(\vec x_1,\vec x_2,\vec x_3,\vec x_4\) are linearly independent.
There is a nonzero vector orthogonal to all four vectors.
Every subset containing exactly three of the four vectors is a basis for \(\mathbb R^3\).
Solution
Every subset containing exactly three of the four vectors is a basis for \(\mathbb R^3\).
Select options 1, 3, and 4.
Impossible. If both additional vectors were on \(P\), all four vectors would span only \(P\), not \(\mathbb R^3\).
Possible. Choose \(\vec x_3\) outside \(P\) and \(\vec x_4=\vec x_1\). These two vectors are independent, and all four span \(\mathbb R^3\).
Impossible. Four vectors in \(\mathbb R^3\) are always linearly dependent.
Impossible. A vector orthogonal to all four is orthogonal to their entire span, \(\mathbb R^3\). In particular, it is orthogonal to itself and must be zero.
Possible. Choose \(\vec x_3\) outside \(P\) and \(\vec x_4=\vec x_1+\vec x_2+\vec x_3\). The first three vectors form a basis. Replacing any one with their sum still gives a basis, since the omitted vector can be recovered by subtracting the other two from the sum.
Problem 7 5 pts
Consider the line \(L\) in \(\mathbb R^3\) defined below.
As \(t\) ranges over \(\mathbb R\), so does \(t+2\). Hence \(L\) is the span of \(\begin{bmatrix}3\\1\\-2\end{bmatrix}\), so it is a subspace of \(\mathbb R^3\).
Describe the set of points that lie on \(L\) and on both planes.
There are no points in the intersection.
A single point.
An entire line of points.
An entire plane of points.
Solution
An entire plane of points.
A single point. A point on \(L\) has coordinates \(x=6+3t\), \(y=2+t\), and \(z=-4-2t\). Substituting into the two plane equations gives
$$ 6+3t=0,\qquad 28+14t=0 $$
Both hold exactly when \(t=-2\), which gives the origin. Thus, the common intersection is \(\lbrace \vec 0\rbrace \).
Problem 8 11 pts
Consider the vectors \(\vec v_1, \vec v_2, \ldots, \vec v_d\in\mathbb R^n\), where \(d\geq 2\). For some scalar \(t\in\mathbb R\), define:
$$ \vec s=t(\vec v_1+\cdots+\vec v_d) $$
Then, for each \(i=1, 2, \ldots,d\), define
$$ \vec u_i = \vec v_i - \vec s $$
a)
5 pts For this part only, suppose \(d=3\). Find an expression for \(\vec u_1 + \vec u_2 + \vec u_3\) in terms of \(t\), \(\vec v_1\), \(\vec v_2\), and \(\vec v_3\). Note that your answer cannot involve \(\vec s\).
Solution
For \(d=3\), \(\vec s=t(\vec v_1+\vec v_2+\vec v_3)\). Therefore,
4 pts Find a value of \(t\) for which \(\vec u_1,\ldots,\vec u_d\) are guaranteed to be linearly dependent, regardless of the original vectors. Give your answer as an expression in terms of \(n\), \(d\), and/or constants.
Solution
As in part a), observe what happens when we sum all of the \(\vec u_i\). This is a linear combination of the \(\vec u_i\), with every coefficient equal to \(1\):
For this value of \(t\), the sum is \(\vec0\). Since the coefficients in this linear combination are all \(1\), they are not all zero. By the definition of linear independence, the \(\vec u_i\) are therefore linearly dependent.
c)
2 pts Suppose \(t>0\). Fill in the \(\boxed{???}\) with the relationship that is guaranteed:
The guaranteed relationship is \(\boxed{\geq}\). Equality can occur, for example, when all the original vectors are the same nonzero vector, so a strict inequality is not guaranteed.