I am learning linear algebra for a machine learning class and have a question about matrix multiplication. The product of two matrices is undefined whenever the rows of the first matrix (reading right to left) do not match the column of the second matrix.
However, say that I would need (for some reason) to multiply a 3x3 matrix by a 2x2 one. Couldn't I just complete the operation by adding the "missing row" with coordinates [0,0]? I am thinking about it because if matrices represent linear transformations of space, then a 2x2 matrix represent a two dimensional transformation. However, isn't a two dimensional transformation simply a transformation where every other dimension is equal to 0?
To illustrate that, let's say I want to apply a linear transformation [-1,0;0,1] to vector [3,3]. The resulting vector would be a two dimensional vector [-3,3]. Now let's say that after this transformation I want to apply another transformation to the same vector, but this time in three dimensions. To keep the example as simple as possible let's use the identity transformation for this: [1,0,0;0,1,0;0,0,1]. Doing this would require to multiply the 3x3 matrix [1,0,0;0,1,0;0,0,1] by the 2x3 one [-1,0;0,1], which is technically not possible. However, if I apply the method above (i.e. adding a third row with all 0 to the 2x3 matrix) I would still be able to compute the transformation and get the result [-1,0,0;0,1,0]. I can then multiply this to my original vector and get [-3,3,0]. The only difference I can see between [-3,3,0] and [-3,3] is that in the first one I am just "explicitly showing" the third dimension as having coordinate 0, whilst in the second I am keeping this implicit.
What am I missing here?
Regards,
Federico
3 Answers
You need to be a little more explicit about what you want to do. Technically speaking, you simply cannot multiply a 2x2 matrix by a 3x3 one. So I'm not sure what you're after there.
But if $A\in\mathcal{M}_{2\times 2}$ and you want to find $B\in\mathcal{M}_{3\times 3}$ so that whenever $(x_1,x_2,x_3)\in\mathbb{R}^3$ we have $A(x_1,x_2)^T$ identical to the projection of $B(x_1,x_2,x_3)^T$ onto its first two coordinates, then that's definitely doable. Indeed, suppose $$ A=\begin{pmatrix} a & b \\ c & d \end{pmatrix} $$ Then you need to let $$ B=\begin{pmatrix} a & b & 0 \\ c & d & 0 \\ * & * & * \end{pmatrix} $$ where the stars can be anything you like.
I'd say the core of the idea can be used, but a matrix containing only 0 doesn't leave the vector unchanged, the unit map does not change the vector, so if you want to use a 2D map on a 3D object you can opt to choose 1 axis along which nothing will be changed using the unit map rather than the zero map.
You're effectively doing calculations in $\mathit{FSM}(\mathbb R)$, the rng of finitely supported $\mathbb N\times\mathbb N$-dimensional real-valued matrices. And you're considering each such matrix to be equivalent to any finite-dimensional slice that contains all of its nonzero entries. This generally works fine: for example, multiplication in $\mathit{FSM}(\mathbb R)$ is associative. But there is no identity matrix, since it would need to have infinitely many $1$s, and that's inconvenient.
That said, just because you can do this doesn't necessarily mean you should! To phrase this in an applied setting: Let's say you're designing a linear algebra package from scratch. If the user tries to multiply two matrices with unexpected shapes, then it's likely a mistake, so it's better to signal a big fat error, rather than silently convert half of their data into zeroes. If the user really wants to perform the multiplication, require them to explicitly describe how they want the inputs to be beaten into shape first. For example, see numpy's array manipulation and pad functions.