Fun with shear operations and SVD – V – matrices of sheared n-dimensional ellipsoids

Posted on 10. October 2023 by eremo

In my previous post of the series

Fun with shear operations and SVD – I – shear matrices and examples created with Blender
Fun with shear operations and SVD – II – Shearing of rectangles and cubes with Python and Matplotlib
Fun with shear operations and SVD – III – Shearing of circles
Fun with shear operations and SVD – IV – Shearing of ellipses

we have studied the transformation of an ellipse by a shear operation. The coordinates of points on an ellipse and the components of respective position vectors fulfill a quadratic equation (quadratic form):

\[
\alpha_o\,x_o^2 \, + \, \beta_o \, x_o y_o \, + \, \gamma_o \, y_o^2 \:=\: \delta_o
\]

An equivalent matrix equation for respective vectors \( \left(\,x_o,\, y_o\,\right)^T \) is

\[
\left(\,x_o,\, y_o\,\right) \,\circ\, \pmb{\operatorname{A}}_q^O \,\circ\, \left(\,x_o,\, y_o\,\right)^T \: = \: \delta_o
\]

The superscript “T” symbolizes the transposition operation. The symmetric (2×2)-matrix \( \pmb{\operatorname{A}}_q^O \) defines the original, unsheared ellipse. The suffix “q” indicates the quadratic form. I have shown how the shear parameter λ_S impacts the coefficients of a corresponding (2×2)-matrix \(\pmb{\operatorname{A}}_q^S \) that defines the sheared ellipse.

What I have not done in the last post is to show how our matrix \(\pmb{\operatorname{A}}_q^S \) is related to a shear matrix \(\pmb{\operatorname{M}}_{sh} \) (see the first post), which describes the effect of the shear on the vectors \( \left(\,x_o,\, y_o\,\right)^T \). I am going to discuss this below. The given matrix relations will also be valid for general n-dimensional ellipsoids.

Matrix relations as discussed below are helpful to accelerate numerical calculations as Numpy (in cooperation with libraries for your OS) provides highly optimized modules for matrix operations. n-dimensional ellipsoids furthermore characterize hyper-surfaces of multivariate normal distributions which appear in certain areas of Machine Learning and respective data.

Matrix describing a centered n-dimensional ellipsoid

We consider n-dimensional and centered ellipsoids whose symmetry centers coincide with the origin of the Euclidean coordinate system [ECS] we work with. A position vector \(\left(\,x_1^o,\, x_2^o\, \cdots x_n^o \right)^T \)

\[
\pmb{x_o} \: = \: \begin{pmatrix} x_1^o \\ x_2^o \\ \vdots \\ x_n^o \end{pmatrix}
\]

is a vector drawn from the origin to a point on the ellipsoid’s hyper-surface. Note that a general vector of a vector space has no reference to a coordinate system’s origin. Therefore the distinction. A general ellipsoid is defined by a quadratic form in the components of its position vectors. The quadratic form is equivalent to the following matrix equation

\[
\left(\pmb{x_o}\right)^T \,\circ\, \pmb{\operatorname{A}}_{qn}^O \,\circ\, \pmb{x_o} \: = \: 1
\]

where \(\pmb{\operatorname{A}}_{qn}^O \) now represents a symmetric (nxn)-matrix. The “\( \circ \)” symbolizes a matrix product.

Note: A coefficient \( \delta \gt 0 \) which we have used in previous posts on the right side of the equation can be included in the coefficient values of the matrix).

Note that the equations above define an ellipsoid up to a translation vector. This is reflected in the fact that the above equation does not create any linear terms.

Quadratic forms not only define ellipsoids. For an ellipsoid we have to assume that the determinant of\(\pmb{\operatorname{A}}_{qn}^O \) is > 0 and that the matrix is invertible:

\[
\operatorname{det} \left(\pmb{\operatorname{A}}_{qn}^O \right) \: \gt \: 0
\]

Note that you could choose an ECS in which the ellipsoid’s principal axes would align with the ECS’s coordinate axes. Such a choice would correspond to a PCA-transformation of the vector data. \(\pmb{\operatorname{A}}_{qn}^O \) would then become diagonal. This corresponds to the fact that a symmetric matrix always has an eigenvalue-decomposition.

Equation of the quadratic form for the sheared ellipsoid

In the first post of this series I have defined a (invertible) shear matrix as a unipotent matrix \( \pmb{\operatorname{M}}_{sh} \) with all coefficients of the lower triangular part, off the diagonal, being equal to 0.0 and all elements on the diagonal being equal to 1:

\[ \pmb{\operatorname{M}}_{sh} \, = \,
\begin{pmatrix}
1 & m_{12} &\cdots & m_{1n}(\ne0) \\
0 & 1 &\cdots & m_{2n} \\
\vdots &\vdots &\ddots &\vdots \\
0 & 0 &\cdots & 1 \end{pmatrix}
\]

Note:

\[ \operatorname{det} \left(\pmb{\operatorname{M}}_{sh}\right) : = \: 1
\]

So an inverse matrix \( \pmb{\operatorname{M}}_{sh}^{-1} \) exists. Shearing our original ellipse (with position vectors \( \pmb{x_o} \)) leads to new vectors \( \pmb{x_S} \):

\[
\pmb{x_S} \: =\: \,\pmb{\operatorname{M}}_{sh} \,\circ\, \pmb{x_o}
\]

We insert \( \pmb{x_S} \) into our defining equation of the original ellipsoid to derive a matrix equation for the sheared ellipsoid:

\[
\left[ \, \pmb{\operatorname{M}}_{sh}^{-1} \,\circ\, \pmb{x_S} \, \right]^T \,\circ\, \pmb{\operatorname{A}}_{qn}^O \,\circ\, \left[ \, \pmb{\operatorname{M}}_{sh}^{-1} \,\circ\, \pmb{x_S} \, \right] \: = \: 1
\]

Giving:

\[
\left(\pmb{x_S}\right)^T \,\circ\, \left[ \, \left( \, \pmb{\operatorname{M}}_{sh}^{-1} \, \right)^T \,\circ\, \pmb{\operatorname{A}}_{qn}^O \,\circ\, \, \pmb{\operatorname{M}}_{sh}^{-1} \, \right] \,\circ\, \pmb{x_S} \: = \: 1
\]

This, obviously, is a new definition equation for a quadratic form in the components of \( \pmb{x_S} \) with a matrix

\[
\pmb{\operatorname{A}}_{qn}^S \:=\: \left( \, \pmb{\operatorname{M}}_{sh}^{-1} \, \right)^T \,\circ\, \pmb{\operatorname{A}}_{qn}^O \,\circ\, \, \pmb{\operatorname{M}}_{sh}^{-1}
\]

We also find:

\[
\operatorname{det}\left(\pmb{\operatorname{A}}_{qn}^S \right) \:=\: \operatorname{det}\left(\pmb{\operatorname{A}}_{qn}^O \right) \: \gt \: 0
\]

From this we can conclude with confidence that we again have gotten a n-dimensional ellipsoid.

Inclusion of a SVD eigendecomposition of M_S

A “Singular Value Decomposition” [SVD] can be applied to any (nxm)-matrix Q (with n > m):

\[
\pmb{\operatorname{Q}} \:=\: \pmb{\operatorname{U}} \,\circ\, \pmb{\operatorname{\Sigma}} \,\circ\, \, \pmb{\operatorname{V}}^T
\]

The (nxn)-matrix U and the (mxm)-matrix V are orthonormal matrices:

\[ \begin{align}
\pmb{\operatorname{U}} \,\circ\, \pmb{\operatorname{U}}^T \:&=\: 1 \\
\pmb{\operatorname{V}} \,\circ\, \pmb{\operatorname{V}}^T \:&=\: 1
\end{align}
\]

Σ is a diagonal (nxm)-matrix with singular values. The column-vectors of U and V are orthogonal singular vectors. Geometrically, U and V can be interpreted s rotational operations.

Therefore, we can decompose a (nxn) upper triangular shear-matrix into two orthonormal (nxn)-matrices U and V plus a diagonal matrix Σ :

\[
\pmb{\operatorname{M}}_{sh} \:=\: \pmb{\operatorname{U}} \,\circ\, \pmb{\operatorname{\Sigma}} \,\circ\, \, \pmb{\operatorname{V}}^T
\]

This leads to

\[ \begin{align}
\pmb{\operatorname{M}}_{sh}^{-1} \:&=\: \left[\pmb{\operatorname{V}}^T\right]^{-1} \,\circ\, \pmb{\operatorname{\Sigma}}^{-1} \,\circ\, \, \pmb{\operatorname{U}}^{-1} \\
&=\: \pmb{\operatorname{V}} \,\circ\, \pmb{\operatorname{\Sigma}}^{-1} \,\circ\, \pmb{\operatorname{U}}^T
\end{align}
\]

This gives us an alternative form to define the inverse shear matrix. Note that the order of the matrices in the matrix products is essential.

An example for the case of a sheared ellipse

We use the example of a sheared ellipse discussed in the last post to verify the results above numerically for a 2-dim case. To write a respective Python/Numpy-program is simple. I will just give you my numerical results below.

We have used an ellipse with the longer and shorter primary axes having values a = 2 and b = 1, respectively. The ellipse was rotated by 60° against the ECS-axes.

The respective (2×2)-matrix \( \pmb{\operatorname{A}}_q^O \) had the following coefficients

\[
\pmb{\operatorname{A}}_q^O \: = \: \begin{pmatrix} \alpha_o & 1/2\, \beta_o \\ 1/2\,\beta_o & \gamma_o \end{pmatrix} \:=\:
\begin{pmatrix} 3.25 & -1.29903811 \\ -1.29903811 & 1.75 \end{pmatrix}
\]

to fulfill

\[
\alpha_o\,x_o^2 \, + \, \beta_o \, x_o y_o \, + \, \gamma_o \, y_o^2 \:=\: \delta_o \:=\: 4.0
\]

The shear matrix (with λ_S = 0.6) was

\[
\pmb{\operatorname{M}}_{sh} \:=\: \begin{pmatrix} 1.0 & 0.6 \\ 0.0 & 1.0 \end{pmatrix}
\]

The resulting sheared ellipse became

For an ellipse we have shown that \(\pmb{\operatorname{A}}_q^S \) is given by

\[
\pmb{\operatorname{A}}_q^S \: = \: \begin{pmatrix} \alpha_o & 1/2\,\left(\beta_o \,-\,2 \alpha_o \lambda_S \right) \\
1/2\,\left(\beta_o \,-\,2 a_o \lambda_S \right) & \alpha_o \, \lambda_S^2 \, -\, \beta_o \lambda_S \,+\, \gamma_o \end{pmatrix}
\]

From this we get the following numerical values:

A_q^S = 
 [[ 3.25       -3.24903811]
 [-3.24903811  4.47884573]]

Via the Python-statement

M_sh_inv = np.linalg.inv(M_sh)

and

A_q^S_2 = M_sh_inv.T @ A_q @ M_sh_inv

we get the following values

A_q^S_2 = 
 [[ 3.25       -3.24903811]
 [-3.24903811  4.47884573]]

Identical! Using

U_sh, S, Vt_sh = np.linalg.svd(M_sh, full_matrices=True)
S_sh = np.diag(S)
M_sh_2 = U_sh @ S_sh @ Vt_sh
M_sh_inv_2 = Vt_sh.T @ np.linalg.inv(S_sh) @ U_sh.T
A_q^S_2 = M_sh_inv_2.T @ A_q @ M_sh_inv_2

we also reproduce the exactly same values.

Conclusion

We have shown how a shear matrix \( \pmb{\operatorname{M}}_{sh} \) transforms the matrix \(\pmb{\operatorname{A}}_{qn}^O \) which defines an un-sheared n-dimensional ellipsoid into a matrix \(\pmb{\operatorname{A}}_{qn}^S \) defining its sheared counterpart. We have also had a glimpse on a SVD decomposition of a shear matrix. The results will enable us in the next post to apply shear operations on a concrete example of a 3-dimensional ellipsoid.

Fun with shear operations and SVD – IV – Shearing of ellipses

Posted on 2. October 2023 by eremo

In the previous posts of this series we got acquainted with shear operations:

Already established results for shearing a circle

Post III focused on the shearing of a circle, which was centered in the Euclidean coordinate system [ECS] we worked with. The shear operation resulted in an ellipse with an inclination against the coordinate axes of our ECS. This was interesting regarding four points:

A circle, which is centered in a chosen ECS, exhibits a continuous rotational symmetry (isotropy). This obviously allows for a decomposition of a shear operation into a sequence of two affine operations in the chosen ECS: a scaling operation (with different factors along the coordinate axes) followed by a rotation (or the other way round). Equivalently: We could switch to another specific ECS which is already rotated by a proper angle against our originally chosen ECS and just perform a scaling operation there.
The rotation angle is determined by the shear parameter λ.
This seems to stand in some contrast to the shearing of figures with only discrete rotational symmetries: We saw for rectangles and cubes that an additional rotation was required to replace the shear operation by a sequence of scaling and rotation operations.
Points (x, y) of circles and ellipses are described by quadratic forms in two dimensions (with some real coefficients α, β, γ, δ):
\[
\alpha \,x^2 \, + \, \beta \, x \, y \, + \, \gamma \, y^2 \:=\: \delta
\]

Quadratic forms play a general role in the mathematical description of cone-sections. (Ellipses are the results of specific cone-sections.)
Ellipses also result from projections of multi-dimensional ellipsoids onto two-dimensional coordinate planes. Multi-dimensional ellipsoids are described by quadratic forms in an ECS covering the ℝⁿ.
Hyper-surfaces for constant probability density values of multivariate normal vector distributions form multi-dimensional ellipsoids. Here we have a link to Machine Learning where key properties of certain objects are often ruled by Gaussian distributions.

From the first point we may expect that a shear operation applied to a multi-dimensional sphere will result in a multi-dimensional ellipsoid – and that such an operation could be replaced by scaling the original sphere (with different factors along n coordinate axes of a n-dimensional ECS) followed by a rotation (or vice versa). We will explicitly investigate this for a 3-dimensional sphere in the next post.

If our assumption were true we would get a first glimpse of the fact that a general multivariate standard distribution can be created by applying a sequence of distinct affine (i.e. linear) operations to a spherical probability distribution. This is discussed in detail in another post-series in this blog.

What is a bit confusing at the moment is that a replacement of a shear operation by simpler affine operations in general seems to require at least two rotations, but only one when we work with centered isotropic bodies. We come back to this point when we discuss the decomposition of a shear matrix by the so called SVD-procedure.

In the previous post of this series we have used the radius of the circle and the shearing parameter λ to derive analytical expressions for the coordinates of special points with extremal values on our ellipse

Points with maximal and minimal y-coordinate values.
Points with a maximal or minimal distance to the symmetry center of the ellipse. I.e. the end-points of the principal diameters of the ellipse.

From the fact that shearing does not change extremal values along the axis perpendicular to the sharing direction we could easily determine the lengths of the ellipse’s principal axes and the inclination angle of the longer axis with the x-axis of our Euclidean coordinate system [ECS].

What do we have in addition? In another mini-series on ellipses

Properties of ellipses by matrix coefficients – I – Two defining matrices (and two more posts)

I have meanwhile described how the geometry of an ellipse is related to its quadratic form and respective coefficients of a symmetric matrix. I call this matrix A_q. It forces the components of position vectors to fulfill an equation based on a quadratic polynomial. Furthermore A_q‘s eigenvalues and eigenvectors define the lengths of the ellipse’s principal axes and their inclination to the axes of our chosen ECS. The matrix coefficients in addition allow us to determine the coordinates of the points with extremal y-values on an ellipse. We will use these results later in the present post.

Objectives of this post: Shearing of a centered, rotated ellipse

In this post I want to show that shearing a given centered, but rotated original ellipse E_O results in another ellipse E_S with a different inclination angle and different sizes of the principal axes.

In addition we will derive the relations of the shearing parameter λ_S with the coefficients of the symmetric matrix \(\pmb{\operatorname{A}}_q^S \) that defines E_S. I also provide formulas for the dependence of E_S‘s geometrical properties on the shear parameter λ_S.

There are two basic prerequisites:

We must show that the application of a shear transformation to the variables of the quadratic form which describes an ellipse E_O results in another proper quadratic form and a related matrix \(\pmb{\operatorname{A}}_q^S \).
The coefficients of the resulting quadratic form and of \(\pmb{\operatorname{A}}_q^S \) must fulfill a mathematical criterion for an ellipse.

We expect point 1 to be valid because a shear operation is just a linear operation.

To get some exercise we approach our goals by first looking at the simple case of shearing an axis-parallel ellipse before extending our considerations to general ellipses with an inclination angle against the coordinate axes of our chosen ECS.

Continue reading →

Properties of ellipses by matrix coefficients – III – coordinates of points with extremal radii

Posted on 9. September 2023 by eremo

A centered, rotated ellipse can be defined by matrices which operate on position-vectors for points on the ellipse. The topic of this post series is the relation of the coefficients of such matrices to some basic geometrical properties of an ellipse. In the previous posts

Properties of ellipses by matrix coefficients – I – Two defining matrices and eigenvalues
Properties of ellipses by matrix coefficients – II – coordinates of points with extremal y-values

we have found that we can use (at least) two matrix based approaches:

One reflects a combination of two affine operations applied to a unit circle. This approach led us to a non-symmetric matrix, which we called A_E. Its coefficients ((a, b), (c, d)) depend on the lengths of the ellipses’ principal axes and trigonometric functions of its rotation angle.
The second approach is based on coefficients of a quadratic form which describes an ellipse as a special type of a conic section. We got a symmetric matrix, which we called A_q.

We have shown how the coefficients α, β, γ of A_q can be expressed in terms of the coefficients of A_E. Another major result was that the eigenvalues and eigenvectors of A_q completely control the ellipse’s properties.

Furthermore, we have derived equations for the lengths σ₁, σ₂ of the ellipse’s principal axes and the rotation angle by which the major axis is rotated against the x-axis of the Cartesian coordinate system [CCS] we work with.

We have also found equations for the components of the position vectors to those points of the ellipse with maximum y-values.

In this post we determine the components of the vectors to the end-points of the ellipse’s principal axes in terms of the coefficients of A_q. Afterward we shall test our formulas by a Python program and plots for a specific example.

Reduced matrix equation for an ellipse

Our centered, but rotated ellipse is defined by a quadratic form, i.e. by a polynomial equation with quadratic terms in the components x_e and y_e of position vectors to points on the ellipse:

\[
\alpha\,x_e^2 \, + \, \beta \, x_e y_e \, + \, \gamma \, y_e^2 \:=\: 1 \,.
\]

The quadratic polynomial can be formulated as a matrix operation applied to position vectors v_E = (x_E, y_E)^T. With the the quadratic and symmetric matrix A_q

\[ \pmb{\operatorname{A}}_q \:=\:
\begin{pmatrix} \alpha & \beta / 2 \\ \beta / 2 & \gamma \end{pmatrix}
\]

we can rewrite the polynomial equation for the centered ellipse as

\[
\pmb{v}_E^T \circ \pmb{\operatorname{A}}_q \circ \pmb{v}_E \:=\: 1\,, \quad \operatorname{with}\: \pmb{v_E} \,=\, \begin{pmatrix} x_E \\ y_E \end{pmatrix}.
\]

Method 1 to determine the vectors to the principal axes’ end points

My readers have certainly noticed that we have already gathered all required information to solve our task. In the first post of this series we have performed an eigendecomposition of our symmetric matrix A_q. We found that the two eigenvectors of A_q for respective eigenvalues λ₁ and λ₂ point along the principal axes of our rotated ellipse:

\[ \begin{align}
\lambda_1 \: &: \quad \pmb{\xi_1} \:=\: \left(\, {1 \over \beta} \left( (\alpha \,-\, \gamma) \,-\, \left[\, \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2\,\right]^{1/2} \right), \: 1 \, \right)^T \,, \\[8pt]
\lambda_2 \: &: \quad \pmb{\xi_2} \:=\: \left(\, {1 \over \beta} \left( (\alpha \,-\, \gamma) \,+\, \left[\, \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2\,\right]^{1/2} \right), \: 1 \, \right)^T \,.
\end{align}
\]

The T symbolizes a transposition operation. The eigenvalues are related to the A_q-coefficients by the following equations:

\[ \begin{align}
\lambda_1 \:&=\: {1 \over 2} \left(\, \left( \alpha \,+\, \gamma \right) \,-\, \left[ \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2 \,\right]^{1/2} \,\right) \,, \\[8pt]
\lambda_2 \:&=\: {1 \over 2} \left(\, \left( \alpha \,+\, \gamma \right) \,+\, \left[ \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2 \,\right]^{1/2} \,\right) \,.
\end{align}
\]

These eigenvalues correspond to the squares of the lengths of the ellipse’s axes.

\[ \begin{align}
\lambda_1 \:&=\: {1 \over \sigma_1^2} \,, \\[8pt]
\lambda_2 \:&=\: {1 \over \sigma_2^2} \,.
\end{align}
\]

Therefore, we can simply take the components of the normalized vectors

\[ \begin{align}
\lambda_1 \: &: \quad \pmb{\xi_1^n} \:=\: {1 \over \|\pmb{\xi_1}\|}\, \pmb{\xi_1} \,, \\[8pt]
\lambda_2 \: &: \quad \pmb{\xi_2^n} \:=\: {1 \over \|\pmb{\xi_2}\|}\, \pmb{\xi_2}
\end{align}
\]

and multiply them with the square-root of the respective eigenvalues to get the vector components to the end-points of the ellipse’s axes:

\[ \begin{align}
\pmb{\xi_1}^{rmax} \:&=\: \left(\,1\,/\, \sqrt{\lambda_1} \, \right) * \pmb{\xi_1^n} \,, \\[8pt]
\pmb{\xi_2}^{rmax} \:&=\: \left(\,1\,/\, \sqrt{\lambda_2} \, \right) * \pmb{\xi_1^n} \,.
\end{align}
\]

This is trivial regarding the algebraic operations, but results in lengthy (and boring) expressions in terms of the matrix coefficients. So, I skip to write down all the terms. (We do not need it for setting up ordered numerical programs.)

Remember that you could in addition replace (α, β, γ) by coefficients (a, b, c, d) of matrix A_E. See the first post of this series for the formulas. This would, however, produce even longer equation terms.

Equation for points with maximum radius values

We define again some convenience variables:

\[ \begin{align}
a_h \,& =\, {\alpha \over \gamma} \,, \\[8pt]
b_h \,& =\, {1 \over 2 } {\beta \over \gamma} \,, \\[8pt]
d_h \,& =\, {1 \over \gamma} \,, \\[8pt]
g_h \,& =\, a_h \,-\, b_h^2 \,, \\[8pt]
f_h \,& =\, 1 \,+\, b_h^2 \,-\, g_h \phantom{\huge{(}}
\end{align}
\]

and

\[
\xi_h \,=\, { \left[\, 4\,d_h\, g_h\, b_h^2 \,+\, d_h\, f_h^2 \,\right] \over 2\, \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] } \phantom{\Huge{(}} \,,
\]

\[
\eta_h \,=\, { b_h^2 \, d_h^2 \over \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] } \phantom{\Huge{(}} \,.
\]

We find for y_E:

\[ \begin{align}
y_E \:&=\: – b_h \, x_E \, \pm \, \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} \\[8pt]
\:&=\: – b_h \, x_E \, \pm \, \left[\,d_h \,-\, g_h\, x_E^2 \, \right]^{1/2} \phantom{\huge{(}} \,.
\end {align}
\]

We pick the y_E with the positive term in the following steps. (The way for the solution with the negative term in y_E is analogous.) The square of y_E is:

\[
y_E^2 \:=\: – d_h \,+\, \left(b^2 \,-\, g\right)\, x_E^2 \,-\, 2 \, b_h \, x_E \,\left[d \,-\, g\,x_E^2 \right]^{1/2} \,.
\]

To find an extremal value of the radius we differentiate and set the derivative to zero:

\[
{\partial \, \left(y_E^2 \,+\, x_E^2\right) \over \partial \, x_E} \:=\: 0 \: \Rightarrow
\]

\[ \begin{align}
& \left(\, 1\,+\, b_h^2 \,-\, g_h\,\right) \, x_E \,-\, b_h \, \left[\, d_h \,-\, g_h\, x_E^2 \, \right]^{1/2} \\[8pt]
&+\, b_h\,g_h\,x_E^2 \, {1 \over \left[\, d_h \,-\, g_h \, x_E^2 \, \right]^{1/2} } \:=\: 0 ,.
\end{align}
\]

This results in

\[
f_h\, x_E \, \left[ d_h \,-\, g_h\, x_E^2\right]^{1/2} \:=\: b_h\,d_h \,-\, 2\, b_h\,g_h\,x_E^2 \, .
\]

Solution for x_e-values of the end-points of the principal axes

We take the square of both sides and reorder terms to get

\[
\left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right]\, x_E^4 \,-\, \left[\, 4\,d_h\, g_h\, b_h^2 \,+\, d_h\, f_h^2 \,\right]\, x_E^2 \,+\, b_h^2 \, d_h^2 \:=\: 0 \,,
\]

\[
x_E^4 \,-\, { \left[\, 4\,d_h\, g_h\, b_h^2 \,+\, d_h\, f_h^2 \,\right] \over \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] } \, x_E^2 \:=\:
-\, { b_h^2 \, d_h^2 \over \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] }
\]

With

\[ \begin{align}
\xi_h \,&=\, { \left[\, 4\,d_h\, g_h\, b_h^2 \,+\, d_h\, f_h^2 \,\right] \over 2\, \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] } \\
\eta_h \,&=\, { b_h^2 \, d_h^2 \over \left[\, 4\,b_h^2\,g_h^2 \,+\, g_h\,f_h^2\,\right] } \phantom{\Huge{)^A}}
\end{align}
\]

we have

\[
x_E^4 \, -\, 2 \,\xi_h \,E_e^2 \:=\: – \eta_h \,.
\]

With the help of a quadratic supplement we get

\[
\left[\, x_E^2 \, -\, \xi_h \right]^2 \:=\: \xi_h^2 \:-\: \eta_h
\]

and find the solution

\[
x_E \:=\: \pm \, \sqrt{ \xi_h \,\pm\, \sqrt{\, \xi_h^2 \,-\,\eta_h \,} } \,.
\]

A detailed analysis also for the other y_E-expression (see above) leads to further solutions for the coordinates (=vector component values) of points with extremal values for the radii. These are the end-points of the principal axes of the ellipse:

\[ \begin{align}
x_{E1}^{rmax} \:&=\: -\, \sqrt{ \, \xi_h \,-\, \sqrt{\, \xi_h^2 \,-\,\eta \,} } \,, \\[8pt]
y_{E1}^{rmax} \:&=\: +\, \sqrt{ \, d_h \,-\, g_h \, \left(x_{e1}^{rmax}\right)^2 \,} \,-\, b_h \, x_{e1}^{rmax} \,, \\[10pt]
x_{E2}^{rmax} \:&=\: -\, x_{e1}^{rmax} \phantom{\huge{)}} \,, \\[8pt]
y_{E2}^{rmax} \:&=\: -\, y_{e1}^{rmax} \,, \\[10pt]
x_{E3}^{rmax} \:&=\: +\, \sqrt{ \, \xi_h \,+\, \sqrt{\, \xi_h^2 \,-\,\eta \,} } \phantom{\huge{)}} \,, \\[8pt]
y_{E3}^{rmax} \:&=\: +\, \sqrt{ \, d_h \,-\, g_h \, \left(x_{e3}^{rmax}\right)^2 \,} \,-\, b_h \, x_{e3}^{rmax} \,, \\[10pt]
x_{E4}^{rmax} \:&=\: -\, x_{e3}^{rmax} \phantom{\huge{)}} \\[8pt]
y_{E4}^{rmax} \:&=\: -\, y_{e3}^{rmax} \,.
\end {align}
\]

I leave it to the reader to expand the convenience variables into terms containing the original coefficients α, β, γ.

Plots

It is easy to write a Python program, which calculates and plots the data of an ellipse and the special points with extremal values of the radii and extremal values of y_e. The general steps which I followed were:

Step 0: Create 100 points a unit circle. Save the coordinates in Python lists (or Numpy arrays). Use Matplotlib’s plot(x,y)-function to plot the vectors.

Step 1: Create an axis-parallel ellipse with values for the axes ha = 2.0 and hb = 1.0 along the x- and the y-axis of the Cartesian coordinate system [CCS]. Do this by applying a diagonal scaling matrix D_{σ1, σ2} (see the first post of this series).

Step 2: Rotate the ellipse bei π/3 (60 °). Do this by applying a rotation matrix R_π/3 to the position vectors of your ellipse (with the help of Numpy). Alternatively, you can first create the matrices, perform a matrix multiplication and then apply the resulting matrix to the position vectors of your unit circle.

(The limiting lines have been calculated by the formulas given above.)

Step 3: Determine the coefficients of combined matrix A_E = R_π/3 ○ D_{σ1, σ2}

I got for the coefficients ( (a, b), (c, d) ) of A_E :

A_ell = 
[[ 1.         -0.8660254 ]
[ 1.73205081  0.5       ]]

Step 3: Determine the coefficients of the matrix A_q by the formulas given in the first post of this series. I got

A_q = 
[[ 3.25       -1.29903811]
[-1.29903811  1.75      ]]

For δ I got:

delta =  4.0

which is consistent with the length-values of the principal axes.

Step 4: Determine values for the eigenvalues λ₁ and λ₂ from the A_q-coefficients by the formulas given in the first post. Also calculate them by using Numpy’s
eigenvalues, eigenvectors = numpy.linalg.eig(A_q). Theory tells us that these values should be exactly λ₁ = 4 and λ₁ = 1. I got

Eigenvalues from A_q:  lambda_1 = 4. :: lambda_2 = 1.

Step 5: Determine the components of the normalized eigenvectors with the help of numpy.linalg.eig(A_q). I got:

Components of normalized eigenvectors by theoretical formulas from A_q coefficients: 
ev_1_n :  -0.8660254037844386  :  0.5000000000000002
ev_2_n :  0.5000000000000001  :  0.8660254037844385

Eigenvectors from A_q via numpyy.linalg.eig():  
ev_1_num :  0.8660254037844387  :  -0.5000000000000001
ev_2_num :  0.5000000000000001  :  0.8660254037844387

The deviation between ev_1_n and ev_1_num is just due to a difference by -1. This is correct as the eigenvectors are unique only up to a minus-sign in all components.

Step 6: Calculate the sinus of the rotation angle of our ellipse from A_q– and A_q-coefficients. The theoretical value is sin(2 π/3) = sin(2 pi/3) = 0.8660254037844387. I got:

sin(2. * rotation angle) of major axis of the ellipse against the CCS x-axis from A_E coefficients: 
sin_2phi-A_E  =  0.8660254037844388

sin(2. * rotation angle) of major axis of the ellipse against the CCS x-axis from from eigenvectors of A_q:
sin_2phi-ev_A_q =  0.8660254037844387 

sin(2. * rotation angle) of major axis of the ellipse against the CCS x-axis from A_q-coefficients:
sin_2phi-coeff-A_q =  0.8660254037844388

Perfect!

Step 7: Plot the end-points of the normalized eigenvectors of A_q:

Note that in our example case the end-point of the eigenvector along the minor axis must be located exactly on the elliptic curve as the ellipses minor axes has a length of b=1!

Step 8: Calculate the components of the vectors to data-points of the ellipse with maximal absolute y_e-values from the A_q-coefficients given in the previous post. Plot these data-points (here in green color).

Step 9: Calculate the components of the vectors to data-points of the ellipse with maximal values of the radii with the help of the complex formulas presented in this post and plot these points in addition.

Conclusion

In this mini-series of posts we have performed some small mathematical exercises with respect to centered and rotated ellipses. We have calculated basic geometrical properties of such ellipses from the coefficients of matrices which define ellipses in algebraic form. Linear Algebra helped us to understand that the eigenvectors and eigenvalues of a symmetric matrix, whose coefficients stem from a quadratic equation (for a conic section), control both the orientation and the lengths of the ellipse’s axes completely.

This knowledge is useful in some Machine Learning [ML] context where elliptic data appear as projections of multivariate normal distributions. Multivariate Gaussian probability functions control properties of a lot of natural objects. Experience shows that certain types of neural networks may transform such data into multivariate normal distributions in latent spaces. An evaluation of the numerical data coming from such ML-experiments often delivers the coefficients of defining matrices for ellipses.

In my blog I now return to the study of with shearing operations applied to circles, spheres, ellipses and 3-dimensional ellipsoids. Later I will continue with the study of multivariate normal distributions in latent spaces of Autoencoders. For both of these topics the knowledge we have gathered regarding the matrices behind ellipses will help us a lot.

Properties of ellipses by matrix coefficients – II – coordinates of points with extremal y-values

Posted on 4. September 2023 by eremo

In the previous post of this mini-series

Properties of ellipses by matrix coefficients – I – Two defining matrices

I have discussed how the coefficients of two matrices which can be used to define a centered, rotated ellipse can be used to calculate geometrical properties of the ellipse:

The lengths σ₁, σ₂ of the ellipse’s principal axes and the rotation angle by which the major axis is rotated against the x-axis of the Cartesian coordinate system [CCS] we work with.

But there are other properties which are interesting, too. A centered, rotated ellipse has two points with extremal values in their y-coordinates. Can we express the coordinates – or equivalently the components of respective position vectors – in terms of the basic matrix coefficients?

The answer is, of course, yes. This post provides a derivation of respective formulas.

Matrix equation for an ellipse

In the last post we have shown that a centered ellipse is defined by a quadratic form, i.e. by a polynomial equation with quadratic terms in the components x_E and y_E of position vectors for points of the ellipse:

\[
\alpha * x_E^2 \, + \, \beta * x_E * y_E \, + \, \gamma * y_E^2 \:=\: 1 \,.
\]

The quadratic polynomial can be formulated as a matrix operation applied to position vectors v_E of points on an ellipse. With the the quadratic and symmetric matrix A_q

\[ \pmb{\operatorname{A}}_q \:=\:
\begin{pmatrix} \alpha & \beta / 2 \\ \beta / 2 & \gamma \end{pmatrix}
\]

we can rewrite the polynomial equation for the centered ellipse as

\[
\pmb{v}_e^T \circ \pmb{\operatorname{A}}_q \circ \pmb{v}_E \:=\: 1 \,, \quad \operatorname{with}\: \pmb{v_E} \,=\, \begin{pmatrix} x_E \\ y_E \end{pmatrix}.
\]

We know already that (under standard conditions)

\[
\operatorname{det}\left( \pmb{\operatorname{A}}_q \right) \:=\: \left( \, \alpha \, \gamma \,-\, {1\over 4}\, \beta^2\, \right) \: \gt \: 0 \,.
\]

A_q is a symmetric and invertible matrix.

Points of the ellipse with extremal y_E – values

Our first step to get the x_E and y_E coordinates for extremal points is to evaluate the quadratic form with respect to y_E. We define:

\[ \begin{align}
a_h \,& =\, {\alpha \over \gamma} \,, \\[8pt]
b_h \,& =\, {1 \over 2 } {\beta \over \gamma} \,, \\[8pt]
d_h \,& =\, {1 \over \gamma} \,.
\end{align}
\]

We reorder terms of the equation and add a supplemental quadratic term on both sides:

\[
y_E^2 \, +\, b_h\,x_E\,y_E \, +\, b_h^2 \, x_E^2 \:=\: d_h \,-\, a_h \, x_E^2 \,+\, b_h^2 \, x_E^2 \,.
\]

We evaluate the complete quadratic term ob the left side to get

\[
y_E \:=\: – b_h \, x_E \, \pm \, \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} \,.
\]

Let us first focus on the positive of the two alternative terms:

\[
y_E \:=\: – b_h \, x_E \,+\, \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} \,.
\]

We get an extremal y_E by evaluating the derivative with respect to x_E

\[
{ \partial \, y_E \over \partial \, x_E } \:=\: 0 \,.
\]

This means

\[
– b_h \, – \, { \left(a_h \,-\, b_h^2 \right)\,x_E \over \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} } \:=\: 0 \,.
\]

Getting rid of the denominator gives

\[
– b_h \, \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} \, – \, \left(a_h \,-\, b_h^2 \right)\,x_E \:=\: 0
\]

and

\[
x_E \:=\: { – b_h \, \left[\,d_h \,-\, \left( a_h \,-\, b_h^2 \right)\, x_E^2 \, \right]^{1/2} \over \left(a_h \,-\, b_h^2 \right) } \,.
\]

Squaring both sides and reordering leads to

\[
x_E^2 \:=\: { – b_h^2 \, d_h \over a_h \, \left(a_h \,-\, b_h^2 \right) } \,.
\]

Using the old coefficients provides the final result for x_e:

\[
x_E \:=\: \pm \, {1 \over 2} \, \beta \, \left[ { 1 \over \alpha \,\left( \alpha \, \gamma \,-\, {1 \over 4} \beta^2 \right) } \right]^{1/2} \,.
\]

We now follow the alternative solution for y_E (see above). After a calculation of the y_E-values from the derived x_E values, we get the components for the two position vectors to the points with extremal y-values on the ellipse:

\[ \begin{align}
x_{E1}^{max} \:&=\: – \, {1 \over 2} \, \beta \, \left[ { 1 \over \alpha \,\left( \alpha \, \gamma \,-\, {1 \over 4} \beta^2 \right) } \right]^{1/2} \,,
\\[8pt]
y_{E1}^{max} \:&=\: – \, {1 \over 2} \, {\beta \over \gamma} \, x_{E1}^{max} \,+\,
\left[ { 1 \over \gamma} \,-\, \left( {\alpha \over \gamma} \,-\, {1 \over 4} {\beta^2 \over \gamma^2} \right) \, \left(x_{E1}^{max}\right)^2 \right]^{1/2}
\end{align}
\]

and

\[ \begin{align}
x_{E2}^{max} \:&=\: + \, {1 \over 2} \, \beta \, \left[ { 1 \over \alpha \,\left( \alpha \, \gamma \,-\, {1 \over 4} \beta^2 \right) } \right]^{1/2} \,, \\[8pt]
y_{E2}^{max} \:&=\: + \, {1 \over 2} \, {\beta \over \gamma} \, x_{E2}^{max} \,-\,
\left[ {1 \over \gamma} \,-\, \left( {\alpha \over \gamma} \,-\, {1 \over 4} {\beta^2 \over \gamma^2} \right) \, \left(x_{E2}^{max}\right)^2 \right]^{1/2} \,.
\end{align}
\]

Solution in terms of the coefficients of an alternative matrix A_E

In my previous post I have discussed yet another matrix A_E which also can be used to define an ellipse. This matrix summarizes two affine transformations of a centered unit circle: A_E = R_φ ○ D_{σ1, σ2}. D_{σ1, σ2}.

You find the relations between the coefficients (a, b, c, d) of matrix A_E and the coefficients (α, β, γ) of matrix A_q in my previous post. This will allow you to calculate the vectors to the extremal points of an ellipse in terms of the coefficients (a, b, c, d).

Conclusion

In this post we have again used the coefficients of a matrix which defines an ellipse via a quadratic form to get information about a geometrical property.
We can now calculate the components of the position vectors to the two points of an ellipse with extremal y-values as functions of the matrix coefficients.

In the next post

Properties of ellipses by matrix coefficients – III – coordinates of points with extremal radii

I will show you how to calculate the components of the end points of the principal axes of the ellipse with the help of our matrix for a quadratic form. I will also use our theoretical results for plots of some ellipses’ axes and of their extremal points. We will also compare theoretical predictions with numerically evaluated values.

Properties of ellipses by matrix coefficients – I – Two defining matrices and eigenvalues

Posted on 4. September 2023 by eremo

For my two current post series on multivariate normal distributions [MNDs] and on shear operations

Multivariate Normal Distributions – II – random vectors and their covariance matrix

Fun with shear operations and SVD

some geometrical and algebraic properties of ellipses are of interest.

Geometrically, we think of an ellipse in terms of their perpendicular two principal axes, focal points, ellipticity and rotation angles with respect to a coordinate system. As elliptic data appear in many contexts of physics, chaos theory, engineering, optics … ellipses are well studied mathematical objects. So, why a post about ellipses in the Machine Learning section of a blog?

In my present working context ellipses appear as a side result of statistical multivariate normal distributions [MNDs]. The projections of multidimensional contour hyper-surfaces of a MND within the ℝⁿ onto coordinate planes of an Cartesian Coordinate System [CCS] result in 2-dimensional ellipses. These ellipses are typically rotated against the axes of the CCS – and their rotation angles reflect data correlations. The general relations of statistical vector data with projections of multidimensional MNDs are somewhat intricate.

Data produced in numerical experiments, e.g. in a Machine Learning context, most often do not give you the geometrical properties of ellipses, which some theory may have predicted, directly. Instead you may get numerical values of statistical vector distributions which correspond to algebraic coefficients. Examples of such coefficients are e.g. correlation coefficients. These coefficients can often be regarded as elements of a matrix. In case of an underlying MND of your statistical variables these matrices indirectly and approximately describe contour surfaces – namely ellipsoids. Or regarding projections of the data on 2-dim coordinate planes “ellipses”.

Ellipsoids/ellipses can in general be defined by matrices operating on position vectors. In particular: Coefficients of quadratic polynomial expressions used to describe ellipses as conic sections correspond to the coefficients of a matrix operating on position vectors.

So, when I became confronted with multidimensional MNDs and their projections on coordinate planes the following questions became interesting:

How can one derive the lengths σ₁, σ₂ of the perpendicular principal axes of an ellipse from data for the coefficients of a matrix which defines the ellipse by a polynomial expression?
By which formula do the matrix coefficients provide the inclination angle of the ellipse’s primary axes with the x-axis of a chosen coordinate system?

You may have to dig a bit to find correct and reproducible answers in your math books. Regarding the resulting mathematical expressions I have had some bad experiences with ChatGPT. But as a former physicist I take the above questions as a welcome exercise in solving quadratic equations and doing some linear algebra. So, for those of my readers who are a bit interested in elementary math I want to answer the posed questions step by step and indicate how one can derive the respective formulas. The level is moderate – you need some knowledge in trigonometry and/or linear algebra.

Centered ellipses and two related matrices

Below I regard ellipses whose centers coincide with the origin of a chosen CCS. For our present purpose we thus get rid of some boring linear terms in the equations we have to solve. We do not loose much of general validity by this step: Results of an off-center ellipse follow from applying a simple translation operation to the resulting vector data. But I admit: Your (statistical) data must give you some relatively precis information about the center of your ellipse. We assume that this is the case.

Our ellipses can be rotated with respect to a chosen CCS. I.e., their longer principal axes may be inclined by some angle φ towards the x-axis of our CCS.

There are actually two different ways to define a centered ellipse by a matrix:

Alternative 1: We define the (rotated) ellipse by a matrix A_E which results from the (matrix) product of two simpler matrices: A_E = R_φ ○ D_{σ1, σ2}.
D_{σ1, σ2} corresponds to a scaling operation applied to position vectors for points located on a centered unit circle. A_φ describes a subsequent rotation. A_E summarizes these geometrical operations in a compact form.
Alternative 2: We define the (rotated) ellipse by a matrix A_q which combines the x- and y-elements of position vectors in a polynomial equation with quadratic terms in the components (see below). The matrix defines a so called quadratic form. Geometrically interpreted, a quadratic form describes an ellipse as a special case of a conic section. The coefficients of the polynomial and the matrix must, of course, fulfill some particular properties.

While it is relatively simple to derive the matrix elements from known values for σ₁, σ₂ and φ it is a bit harder to derive the ellipse’s properties from the elements of either of the two defining matrices. I will cover both matrices in this post.

For many practical purposes the derivation of central elliptic properties from given elements of A_q is more relevant and thus of special interest in the following discussion.

Matrix A_E of a centered and rotated ellipse: Scaling of a unit circle followed by a rotation

Our starting point is a unit circle C whose center coincides with our CCS’s origin. The components of vectors v_c to points on the circle C fulfill the following conditions:

\[
\pmb{C} \::\: \left\{ \pmb{v}_c \:=\: \begin{pmatrix} x_c \\ y_c \end{pmatrix} \:=\: \begin{pmatrix} \operatorname{cos}(\psi) \\ \operatorname{sin}(\psi) \end{pmatrix}, \quad 0\,\leq\,\psi\, \le 2\pi \right\}
\]

and

\[
x_c^2 \, +\, y_c^2 \,=\; 1 \,.
\]

We define an ellipse E(σ₁, σ₂) by the application of two linear operations to the vectors of the unit circle:

\[
\pmb{E}_{\sigma_1, \, \sigma_2} \::\: \left\{ \, \pmb{v}_E \:=\: \begin{pmatrix} x_E \\ y_E \end{pmatrix} \:=\: \pmb{\operatorname{R}}_{\phi} \circ \pmb{\operatorname{D}}_E \circ \pmb{v}_c , \quad \pmb{v_c} \in \pmb{C} \, \right\} \,.
\]

D_E is a diagonal matrix which describes a stretching of the circle along the CCS-axes, and R_φ is an orthogonal rotation matrix. The stretching (or scaling) of the vector-components is done by

\[
\pmb{\operatorname{D}}_E \:=\: \begin{pmatrix} \sigma_1 & 0 \\ 0 & \sigma_2 \end{pmatrix} \,,
\]

\[
\pmb{\operatorname{D}}_E^{-1} \:=\: \begin{pmatrix} {1 / \sigma_1} & 0 \\ 0 & {1 / \sigma_2} \end{pmatrix} \,.
\]

The coefficients σ₁, σ₂ obviously define the lengths of the principal axes of the yet unrotated ellipse.

To be more precise: &sigma;₂ is half of the diameter in y-direction. I.e. σ₁ and σ₂ refer to the half-axes of the ellipse.

The subsequent rotation by an angle φ against the x-axis of the CCS is done by

\[
\pmb{\operatorname{R}}_{\phi} \:=\:
\begin{pmatrix} \operatorname{cos}(\phi) & – \,\operatorname{sin}(\phi) \\ \operatorname{sin}(\phi) & \operatorname{cos}(\phi)\end{pmatrix}
\:=\: \begin{pmatrix} u_1 & -\,u_2 \\ u_2 & u_1 \end{pmatrix} \,,
\]

\[
\pmb{\operatorname{R}}_{\phi}^T \:=\: \pmb{\operatorname{R}}_{\phi}^{-1} \:=\: \pmb{\operatorname{R}}_{-\,\phi} \,\,.
\]

The combined linear transformation results in a matrix A_E with coefficients ((a, b), (c, d)):

\[ \begin{align}
\pmb{\operatorname{A}}_E \:=\: \pmb{\operatorname{R}}_{\phi} \circ \pmb{\operatorname{D}}_E \:=\:
\begin{pmatrix} \sigma_1\,u_1 & -\,\sigma_2\,u_2 \\ \sigma_1\,u_2 & \sigma_2\,u_1 \end{pmatrix} \:=\:: \begin{pmatrix} a & b \\ c & d \end{pmatrix} \,\,.
\end{align}
\]

These is the first set of matrix coefficients we are interested in.

Note:

\[ \pmb{\operatorname{A}}_E^{-1} \:=\:
\begin{pmatrix} {1 \over \sigma_1} \,u_1 & {1 \over \sigma_1}\,u_2 \\ -{1 \over \sigma_2}\,u_2 & {1 \over \sigma_2}\,u_1 \end{pmatrix} \, ,
\]

\[
\pmb{v}_E \:=\: \begin{pmatrix} x_E \\ y_E \end{pmatrix} \:=\: \pmb{\operatorname{A}}_E \circ \begin{pmatrix} x_c \\ y_c \end{pmatrix} \,,
\]

\[
\pmb{v}_c \:=\: \begin{pmatrix} x_c \\ y_c \end{pmatrix} \:=\: \pmb{\operatorname{A}}_E^{-1} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, .
\]

We use

\[ \begin{align}
&u_1 \,=\, \operatorname{cos}(\phi),\quad u_2 \,=\,\operatorname{sin}(\phi), \quad u_1^2 \,+\, u_2^2 \,=\, 1 \,,\\[8pt]
&(1/\lambda_1) \,: =\, \sigma_1^2, \quad\quad (1/\lambda_2) \,: =\, \sigma_2^2
\end{align}
\]

and find

\[ \begin{align}
a \,&=\, \sigma_1\,u_1, \quad b \,=\, -\, \sigma_2\,u_2 \,, \\[8pt]
c \,&=\, \sigma_1\,u_2, \quad d \,=\, \sigma_2\,u_1 \,,
\end{align}
\]

\[
\operatorname{det}\left( \pmb{\operatorname{A}}_E \right) \:=\: a\,d \,-\, b\,c \:=\: \sigma_1\, \sigma_2 \,.
\]

As we already know, σ₁ and σ₂ are factors which give us the lengths of the principal axes of the ellipse. σ₁ and σ₂ have positive values. We, therefore, can safely assume:

\[
\operatorname{det}\left( \pmb{\operatorname{A}}_E \right) \:=\: a\,d \,-\, b\,c \:\gt\: 0
\]

Ok, we have defined an ellipse via an invertible matrix A_E, whose coefficients are directly based on geometrical properties.

But as said: Often an ellipse is described by a an equation with quadratic terms in x and y coordinates of data points. The quadratic form has its background in algebraic properties of conic sections. As a next step we derive such a quadratic equation and relate the coefficients of the quadratic polynomial with the elements of our matrix A_E. The result will in turn define another very useful matrix A_q.

Quadratic forms – Case 1: Centered ellipse, principal axes aligned with CCS-axes

We start with a simple case. We take a so called axis-parallel ellipse which results from applying only a scaling matrix D_E onto our unit circle C. I.e., in this case, the rotation matrix is assumed to be just the identity matrix. We can omit it from further calculations:

\[
\pmb{\operatorname{R}}_{\phi} \:=\:
\begin{pmatrix} 1 & 0 \\ 0 & 1 \end{pmatrix} \,=\, \pmb{\operatorname{I}}, \quad u_1 \,=\, 1,\: u_2 \,=\, 0, \: \phi = 0 \,.
\]

We need an expression in terms of (x_E, y_E).

To get quadratic terms of vector components it often helps to invoke a scalar product. The scalar product of a vector with itself gives us the squared norm or length of a vector. In our case the norms of the inversely re-scaled vectors obviously have to fulfill:

\[
\left[\, \pmb{\operatorname{D}}_E^{-1} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, \right]^T \,\bullet \, \left[\, \pmb{\operatorname{D}}_E^{-1} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \,\right] \:=\: 1 \,\,.
\]

(The bullet represents the scalar product of the vectors.) This directly results in:

\[
{1 \over \sigma_1^2} * x_E^2 \, + \, {1 \over \sigma_2^2} * y_E^2 \:=\: 1 \,.
\]

We eliminate the denominator to get a convenient quadratic form:

\[
\lambda_1 * x_E^2 \,+\, \lambda_2 * y_E^2 \:=1 \,.
\]

If we were given the quadratic form more generally by coefficients α, β and γ

\[
\alpha * x_E^2 \,+\, \beta * x_E \, y_E \,+\, \gamma * y_E^2 \,:=\, 1 \,,
\]

we could directly relate these coefficients with the geometrical properties of our ellipse:

Axis-parallel ellipse:

\[ \begin{align}
\alpha \,&=\, 1 / a^2 \,=\, 1 / \sigma_1^2 \,=\, \lambda_1 \,, \\[8pt]
\gamma \,&=\, 1 / d^2 \,=\, 1 / \sigma_2^2 \,=\, \lambda_2 \,, \\[8pt]
\beta \,&=\, b \,=\, c \,=\, 0 \,, \\[8pt]
\phi &= 0 \,.
\end{align}
\]

I.e., we can directly derive σ₁, σ₂ and φ from the coefficients of the quadratic form. But an axis-parallel ellipse is a very simple ellipse. Things get more difficult for a rotated ellipse.

Quadratic forms – Case 2: General centered and rotated ellipse

We perform the same trick with the vectors v_E to get a quadratic polynomial for a rotated ellipse:

\[
\left[ \,\pmb{\operatorname{A}}_E^{-1} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, \right]^T \, \bullet \,
\left[ \, \pmb{\operatorname{A}}_E^{-1} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \,\right] \:=\: 1 \,\,.
\]

I skip the lengthy, but simple algebraic calculation. We get (with our matrix elements a, b, c, d):

\[
{1 \over \sigma_1^2 \, \sigma_2^2 } \, \left[ \,
\left( c^2 \,+\, d^2 \right) * x_E^2 \,\, – \,\, 2\left( a\,c\, +\, b\,d \right)* x_E \, y_E \,\, + \,\, \left(a^2 \,+\, b^2\right) * y_E^2 \, \right]
\,=\, 1 \, .
\]

The rotation has obviously lead to mixing of components in the polynomial. The coefficient for x_E * y_E is > 0 for the non-trivial case.

Quadratic form: A matrix equation to define an ellipse

We rewrite our equation again with general coefficients α, β and γ

\[
\alpha * x_E^2 \, + \, \beta * x_E \, y_E \, + \, \gamma * y_E^2 \,=\, 1 \,.
\]

These are coefficients which may come from some theory or from averages of numerical data.

The quadratic polynomial can in turn be reformulated as a matrix operation with a symmetric matrix A_q:

\[
\pmb{v}_E^T \circ \pmb{\operatorname{A}}_q \circ \pmb{v}_E \:=\: 1 \,,
\]

with

\[ \pmb{\operatorname{A}}_q \:=\:
\begin{pmatrix} \alpha & \beta / 2 \\ \beta / 2 & \gamma \end{pmatrix} \,\,,
\]

\[ \pmb{\operatorname{A}}_q \:=\: { 1 \over \sigma_1^2 \, \sigma_2^2} \,
\begin{pmatrix} c^2 \,+\, d^2 & a\,c \, +\, b\,d \\ a\,c \, +\, b\,d & a^2 \,+\, b^2 \end{pmatrix} \, .
\]

This means

\[ \begin{align}
\alpha \:&=\: { 1 \over \sigma_1^2 \, \sigma_2^2} \, \left(\, c^2 \,+\, d^2 \,\right) \:=\: \lambda_2 * u_2^2 \, + \, \lambda_1 * u_1^2 \,, \\[8pt]
\gamma \:&=\: { 1 \over \sigma_1^2 \, \sigma_2^2} \, \left(\, a^2 \,+\, b^2 \,\right) \:=\: \lambda_2 * u_1^2 \, + \, \lambda_1 * u_2^2 \,, \\[8pt]
\beta \:&=\: – \, 2\, { 1 \over \sigma_1^2 \, \sigma_2^2} \, \left(\,a \,c \,+\, b \, d \,\right) \:=\: – \, 2 \, \left( \lambda_2 \,-\, \lambda_1 \right)\, u_1 \, u_2 \,.
\end{align}
\]

With

\[ u_1 = \cos \phi \,, \quad u_2 = \sin \phi \,,
\]

it follows that

\[ \begin{align}
\alpha \:&=\: \lambda_2 * \sin^2 \phi \, + \, \lambda_1 * \cos^2 \phi \,, \\[8pt]
\gamma \:&=\: \lambda_2 * \cos^2 \phi \, + \, \lambda_1 * \sin^2 \phi \,, \\[8pt]
\beta \:&=\: – \, 2 \, \left( \lambda_2 \,-\, \lambda_1 \right)\, \cos \phi \, \sin \phi \,.
\end{align}
\]

These terms are intimately related to the geometrical data; expect them to play a major role in further considerations.

With the help of the coefficients of A_E we can also show that det(A_q) > 0:

\[
\operatorname{det}\left( \pmb{\operatorname{A}}_q \right) \:=\: \left(\,\alpha \, \gamma \,-\, {1\over 4}\, \beta^2 \, \right) \:=\: { 1 \over \sigma_1^2 \, \sigma_2^2} \, \left(\, b\,c \,-\, a\,d \, \right)^2 \, \gt \, 0 \,.
\]

Thus A_q is an invertible matrix if A_E is invertible. For standard conditions (σ₁ >0, σ₂ > 0) this is the case (see above). Furthermore, A_q is symmetric and thus its own transposed matrix.

Above we have got α, β, γ as some relatively simple functions of a, b, c, d. The inversion is not so trivial and we do not even try it here.

Instead we focus on how we can express σ₁, σ₂ and φ as functions of either (a, b, c, d) or (α, β, γ).

How to derive σ₁, σ₂ and φ from the coefficients of A_E or A_q in the general case?

Let us assume we have (numerical) data for the coefficients of the quadratic form. Then we may want to calculate values for the length of the principal axes and the rotation angle φ of the corresponding ellipse. There are two ways to derive respective formulas:

Approach 1: Use trigonometric relations to directly solve the equation system.
Approach 2: Use an eigenvector decomposition of A_q.

Both ways are fun!

Direct derivation of σ₁, σ₂ and φ from A_q by using trigonometric relations

Trigonometric relations which we can use are:

\[ \begin{align}
\sin (2 \, \phi) \:&=\: 2 \,\cos (\phi)\, \sin(\phi) \,, \\[8pt]
\cos (2 \, \phi) \:&=\: 2 \,\cos^2 (\phi)\, -\, 1 \\[8pt]
\:&=\: 1 \,-\, 2\,\sin^2(\phi) \\[8pt]
\:&=\: \cos^2 (\phi)\, -\, \sin^2 (\phi) \,.
\end{align}
\]

Without loosing much of generality we further assume

\[
\lambda_2 \:\ge \lambda_1 \,.
\]

This affects some aspects of the following derivations and should be kept in mind whilst reading. In the end, our results would only differ by a rotation of π/2, if we had chosen otherwise.

Note, however, that due to our assumption we discuss ellipses whose half-axis σ₁ in x-direction is longer than the half-axis σ₂ in y-direction!

By combining the above relations for the A_q-coefficients we find

\[ \begin{align}
\alpha \, + \, \gamma \,&=\, \lambda_1 \,+\, \lambda_2 \,, \\[8pt]
\gamma \,\, – \, \alpha \,&=\, \left( \, \lambda_2 \,-\, \lambda_1\,\right) \, \cos (2\,\phi) \,, \\[8pt]
\,-\, \beta \:&=\: \left( \, \lambda_2 \,-\, \lambda_1 \, \right)\, \sin (2\,\phi) \,.
\end{align}
\]

By squaring and adding the last two equations we further get

\[ \begin{align}
\lambda_2 \,+\, \lambda_1 \,&=\, \alpha \, + \gamma \,, \\[8pt]
\lambda_2 \,-\, \lambda_1 \,&=\, \left[\, \beta^2 \,+\, \left(\,\gamma \,-\,\alpha \,\right)^2\, \right]^{1/2} \,.
\end{align}
\]

Eventually we conclude:

\[ \begin{align}
\lambda_1 \,&=\, {1\over 2}\, \left[\, (\,\alpha\,+\,\gamma\,) \,-\, \left[\, \beta^2 \,+\, \left(\,\gamma \,-\,\alpha \,\right)^2\, \right]^{1/2} \, \right] \,, \\[8pt]
\lambda_2 \,&=\, {1\over 2}\, \left[\, (\,\alpha\,+\,\gamma\,) \,+\, \left[\, \beta^2 \,+\, \left(\,\gamma \,-\,\alpha \,\right)^2\, \right]^{1/2} \, \right] \,.
\end{align}
\]

Note that λ₁ indeed is smaller than λ₂! This reflects our assumption made above about the lengths of the ellipse’s axes.

For λ₂ ≥ λ₁ the rotation angle is given by

\[
\phi \:=\: \, {1 \over 2} \operatorname{arcsin}\left( {-\, \beta \over \left[ \beta^2 \, +\, \left( \gamma \,-\, \alpha \right)^2 \, \right]^{1/2} } \right) \,.
\]

See a later section below for ambiguities coming from the arcsin-function. A proper analysis and interpretation of this formula for the rotation angle is necessary, when you want to reconstruct ellipses by numerical methods from a given matrix A_q. Some matrices may describe ellipses with their half-axis in y-direction being longer than the half-axis in x-direction. Then λ₁ and λ₂ would change their role with respect to x,y.

Example:

A typical example for a matrix would be one that appears in the context of a standardized (!) bivariate normal distribution with some correlation imposed onto the statistical variables:

\[ \pmb{\operatorname{A}}_q \:=\: {1\over 1\,-\, \rho^2}\,
\begin{pmatrix} 1 & -\,\rho \\ -\,\rho & 1 \end{pmatrix} \,\,,
\]

with ρ > 0 being the Pearson correlation coefficient. For this case we get for the half axes of a respective ellipse and the angle :

\[ \begin{align}
\sigma_1 \,& =\, \sqrt{\left(\, 1\,+\, \rho\,\right)\, } \,, \\[8pt]
\sigma_2 \,& =\, \sqrt{\left(\, 1\,-\, \rho\,\right)\, } \,, \\[8pt]
\phi \,& =\, \pi / 4 \,. \\[8pt]
\end{align}
\]

Direct derivation of σ₁, σ₂ and φ from A_E by using trigonometric relations

We set

\[ \begin{align}
\epsilon_1 \,&=\, \sigma_1^2 \,=\, {1 \over \lambda_1 } \,, \\[8pt]
\epsilon_2 \,&=\, \sigma_2^2 \,=\, {1 \over \lambda_2 } \,.
\end{align}
\]

We have:

\[ \begin{align}
a^2 \,+\, b^2 \:&=\: \epsilon_1 * \cos^2 \phi \, + \, \epsilon_2 * \sin^2 \phi \,, \\[8pt]
c^2 \,+\, d^2 \:&=\: \epsilon_1 * \sin^2 \phi \, + \, \epsilon_2 * \cos^2 \phi \,, \\[8pt]
a\, c \,+\, b\, d \:&=\: \left( \, \epsilon_1 \,-\, \epsilon_2 \, \right) * \cos \phi * \sin \phi \,.
\end{align}
\]

This leads to

\[ \begin{align}
2 \,\left( \, a^2 \,+\, b^2 \, \right) \:&=\: \epsilon_1 * \left(\, 1 \,+\, \cos (2\phi) \, \right) \, + \, \epsilon_2 * \left( \, 1 \,-\, \cos (2\phi) \, \right) \,, \\[8pt]
2 \,\left( \, c^2 \,+\, d^2 \, \right) \:&=\: \epsilon_2 * \left(\, 1 \,+\, \cos (2\phi) \, \right) \, + \, \epsilon_1 * \left( \, 1 \,-\, \cos (2\phi)\, \right) \,, \\[8pt]
2 \, \left(\,a \,c \,+\, b\, d \,\right) \:&=\: \left(\, \epsilon_1 \,-\, \epsilon_2 \, \right) * \sin (2\phi) \,.
\end{align}
\]

We rearrange terms and get:

\[ \begin{align}
\epsilon_1 \,+\, \epsilon_2 \, +\, \left( \epsilon_1 \,-\, \epsilon_2 \right) * \cos (2\phi) \:&=\: \,+ 2 \, \left( \, a^2 \,+\, b^2 \, \right) \,, \\[8pt]
\,-\, \epsilon_1 \,-\, \epsilon_2 \,+\, \left( \epsilon_1 \,-\, \epsilon_2 \right) * \cos (2\phi) \:&=\: \,- 2 \, \left( \, c^2 \,+\, d^2 \, \right) \,, \\[8pt]
\left(\, \epsilon_1 \,-\, \epsilon_2 \, \right) * \sin (2\phi) \:&=\: \,+ 2 \, \left( \, a\,c \,+\, b\,d \, \right) \,.
\end{align}
\]

Let us define some further variables before we add and subtract the first two of the above equations:

\[ \begin{align}
r \:&=\: {1 \over 2} \left(\, a^2 \,+\, b^2 \,+\, c^2 \,+\, d^2 \, \right) \,, \\[8pt]
s_1 \:&=\: {1 \over 2} \left(\, a^2 \,+\, b^2 \,-\, c^2 \,-\, d^2 \, \right) \,, \\[8pt]
s_2 \:&=\: \left(\, a\,c \,+\, b\,d \, \right) \,, \\[8pt]
\pmb{s} \:&=\: \begin{pmatrix} s_1 \\ s_2 \end{pmatrix} \,, \\[8pt]
s \:&=\: \sqrt{ s_1^2 \,+\, s_2^2 } \, .
\end{align}
\]

Adding two of the equations with the sin(2φ) and cos(2φ) above and using the third equation results in:

\[
{1 \over 2} \, \left( \epsilon_1 \,-\, \epsilon_2 \right) \begin{pmatrix} \cos (2\phi) \\ \sin (2\phi) \end{pmatrix} \:=\: \begin{pmatrix} s_1 \\ s_2 \end{pmatrix} \,.
\]

Taking the vector norm on both sides (with ε₁ ≥ ε₂) and adding two of the equations above results in:

\[ \begin{align}
\epsilon_1 \,\, – \,\, \epsilon_2 \:&=\: 2 * s \,, \\[8pt]
\epsilon_1 \, + \, \epsilon_2 \:&=\: 2 * r \,.
\end{align}
\]

This gives us:

\[ \begin{align}
\sigma_1^2 \,=\, \epsilon_1 \:&=\: r \,+\, s \,, \\[8pt]
\sigma_2^2 \,=\, \epsilon_2 \:&=\: r \,-\, s \,.
\end{align}
\]

In terms of the coefficients a, b, c, d:

\[ \begin{align}
\sigma_1^2 \,=\, {1 \over \lambda_1} \:&=\: {1 \over 2} \left[ \, a^2+b^2+c^2 +d^2 \,+\, \left[ 4 (ac + bd)^2 \, +\, \left( c^2+d^2 -a^2 -b^2\right)^2 \, \right]^{1/2} \right] \,, \\[8pt]
\sigma_2^2 \,=\, {1 \over \lambda_2} \:&=\: {1 \over 2} \left[ \, a^2+b^2+c^2 +d^2 \,-\, \left[ 4 (ac + bd)^2 \, +\, \left( c^2+d^2 -a^2 -b^2\right)^2 \, \right]^{1/2} \right] \,.
\end{align}
\]

Who said that life had to be easy?

But, it is relatively easy to prove:

\[
\sigma_1^2 * \sigma_2^2 \,=\, {1 \over \lambda_1^2 \, \lambda_2^2} \,=\, \left(\, a\,d\,-\, b\,c\ \, \right)^2 \,=\, \left[\operatorname{det}\left(\pmb{\operatorname{A}}_q \right)\right]^2 \,.
\]

So, we can indeed confirm :

\[ \begin{align}
\alpha \:&=\: { 1 \over \left(\,a\,d \,-\, b\,c\,\right)^2 } \, \left(\, c^2 \,+\, d^2 \,\right) \,, \\[8pt]
\gamma \:&=\: { 1 \over \left(\,a\,d \,-\, b\,c\,\right)^2 } \, \left(\, a^2 \,+\, b^2 \,\right) \,, \\[8pt]
\beta \:&=\: – \, 2\, { 1 \over \left(\,a\,d \,-\, b\,c\,\right)^2 } \, \left(\,a \,c \,+\, b \, d \,\right)\,.
\end{align}
\]

Determination of the inclination angle φ

For the determination of the angle φ we use:

\[
\begin{pmatrix} \operatorname{cos}(2\phi) \\ \operatorname{sin}(2\phi) \end{pmatrix} \:=\: {1 \over s} \begin{pmatrix} s_1 \\ s_2 \end{pmatrix} \,.
\]

We choose

\[
-\pi/2 \,\lt\, \phi \le \pi/2 \,,
\]

and get:

\[
\phi \:=\: {1 \over 2} \operatorname{arctan}\left({s_2 \over s_1}\right) \:=\: {1 \over 2} \, \operatorname{arctan}\left( { 2\left(ac \,+\, bd\right) \over (a^2 \,+\, b^2 \,-\, c^2 \,-\, d^2) } \right) \,.
\]

Note: All in all there are four different solutions. The reason is that we alternatively could have requested λ₂ ≥ λ₁ and also chosen the angle π + φ. So, the ambiguity is due to a selection of the considered principal axis and rotational symmetries.

2nd way to a solution for σ₁, σ₂ and φ via eigendecomposition

For our second way of deriving formulas for σ₁, σ₂ and φ we use some linear algebra.

This approach is interesting for two reasons: It indicates how we can use the Python “linalg”-package together with Numpy to get results numerically. In addition we get familiar with a representation of the ellipse in a properly rotated CCS.

Above we have written down a symmetric matrix A_q describing an operation on the position vectors of points on our rotated ellipse:

\[
\pmb{v}_E^T \circ \pmb{\operatorname{A}}_q \circ \pmb{v}_E \:=\: 1 \,.
\]

We know from linear algebra that every symmetric matrix can be decomposed into a product of orthogonal matrices O, O^T and a diagonal matrix. This reflects the so called eigendecomposition of a symmetric matrix. It is a unique decomposition in the sense that it has a uniquely defined solution in terms of the coefficients of the following matrices:

\[
\pmb{\operatorname{A}}_q \:=\: \pmb{\operatorname{O}} \circ \pmb{\operatorname{D}}_{diag} \circ \pmb{\operatorname{O}}^T
\]

with

\[
\pmb{\operatorname{D}}_{diag} \:=\: \begin{pmatrix} \lambda_{u} & 0 \\ 0 & \lambda_{d} \end{pmatrix}
\]

The coefficients λ_u and λ_d are eigenvalues of both D_diag and A_q.

Reason:
Orthogonal matrices do not change eigenvalues of a transformed matrix. So, the diagonal elements of D_diag are the eigenvalues of A_q. Linear algebra also tells us that the columns of the matrix O are given by the components of the normalized eigenvectors of A_q.

We can interpret O as a rotation matrix R_ψ for some angle ψ:

\[
\pmb{v}_E^T \circ \pmb{\operatorname{A}}_q \circ \pmb{v}_E \:=\: \pmb{v}_E^T \circ
\pmb{\operatorname{R}}_{\psi} \circ \pmb{\operatorname{D}}_{diag} \circ \pmb{\operatorname{R}}_{\psi}^T \circ \pmb{v}_E \:=\: 1 \,.
\]

This means

\[
\left[ \pmb{\operatorname{R}}_{-\psi} \circ \pmb{v}_E \right]^T \circ
\pmb{\operatorname{D}}_{diag} \circ \left[ \pmb{\operatorname{R}}_{-\psi} \circ \pmb{v}_E \right] \:=\: 1 \,.
\]

The whole operation tells us a simple truth, which we are already familiar with. By our construction procedure for a rotated ellipse we know that a rotated CCS exists, in which the ellipse can be described as the result of a scaling operation (along the coordinate axes of the rotated CCS) applied to a unit circle. (This CCS is, of course, rotated by an angle φ against our working CCS in which the ellipse appears rotated.)

Above we had found

With our matrices R_φ and scaling matrix D_E we can rewrite this as

\[
\left[ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, \right]^T \, \circ \,
\left[ \pmb{\operatorname{D}}_E^{-1} \circ \pmb{\operatorname{R}}_{-\phi} \right]^T \bullet
\left[ \pmb{\operatorname{D}}_E^{-1} \circ \pmb{\operatorname{R}}_{-\phi} \right] \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \:=\: 1 \,.
\]

So:

\[
\left[ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, \right]^T \, \circ \,
\left[ \pmb{\operatorname{R}}_{-\phi} \right]^T \circ \left[ \pmb{\operatorname{D}}_E^{-1} \right]^T \circ
\pmb{\operatorname{D}}_E^{-1} \circ \pmb{\operatorname{R}}_{-\phi} \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \:=\: 1 \,.
\]

Remembering that a diagonal matrix is its own transposed matrix and that the inverse of an orthogonal matrix (rotation) is its transposed matrix, we get:

\[
\left[ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \, \right]^T \, \circ \,
\left[ \pmb{\operatorname{R}}_{\phi} \right] \circ \left[ \pmb{\operatorname{D}}_E^{-1} \circ
\pmb{\operatorname{D}}_E^{-1} \right] \circ \pmb{\operatorname{R}}_{\phi}^T \circ \begin{pmatrix} x_E \\ y_E \end{pmatrix} \:=\: 1 \,.
\]

Comparing expressions, we find

\[
\pmb{\operatorname{D}}_{diag} \:=\: \left[ \pmb{\operatorname{D}}_E^{-1} \circ
\pmb{\operatorname{D}}_E^{-1} \right] \,.
\]

And thus, following our old assumption λ₁ ≤ λ₂ (!) – we set

\[ \begin{align}
\lambda_u \,=\, \lambda_1 \,&=\, {1 \over \sigma_1^2} \,, \\[8pt]
\lambda_d \,=\, \lambda_2 \,&=\, {1 \over \sigma_2^2} \,.
\end{align} \]

The eigenvalues of our symmetric matrix A_q are just our parameters λ₁ and λ₂.

It is really noteworthy that the half-axes of the ellipse are given by the reciprocate value of the square root of the matrix’ eigenvalues:

\[ \begin{align}
\sigma_1 \:&=\: {1 \over \sqrt{\lambda_1} } \,, \\[8pt]
\sigma_2 \:&=\: {1 \over \sqrt{\lambda_2} } \,.
\end{align} \]

Mathematically, a lengthy calculation (see below) will indeed reveal that the eigenvalues of a symmetric matrix A_q with coefficients α, 1/2*β and γ have the following form:

\[
\lambda_{1/2} \:=\: \lambda_{u/d} \:=\: {1 \over 2} \left[\, \left(\alpha \,+\, \gamma \right) \,\mp\, \left[ \beta^2 + \left(\gamma \,-\, \alpha \right)^2 \,\right]^{1/2} \, \right]
\]

Side remark: You can derive this results a bit easier with solving the so called “characteristic equation” of the matrix. For more details see e.g.:
Eigenvalues and eigenvector of a positive-definite, real valued and symmetric matrix

This is, of course, exactly what we have found some minutes ago by solving respective equations with the help of trigonometric terms. Remember, however, the assumptions about the λ-values and lengths of the ellipse’s axes!

We will prove the fact that these indeed are valid eigenvalues in a minute. Let us first look at respective eigenvectors ξ_1/2. To get them we must solve the equations resulting from

\[
\left( \begin{pmatrix} \alpha & \beta / 2 \\ \beta / 2 & \gamma \end{pmatrix} \,-\, \begin{pmatrix} \lambda_{1/2} & 0 \\ 0 & \lambda_{1/2} \end{pmatrix} \right) \,\circ \, \pmb{\xi_{1/2}} \:=\: \pmb{0},
\]

with

\[
\pmb{\xi_1} \,=\, \begin{pmatrix} \xi_{1,x} \\ \xi_{1,y} \end{pmatrix}, \quad \pmb{\xi_2} \,=\, \begin{pmatrix} \xi_{2,x} \\ \xi_{2,y} \end{pmatrix}
\]

Below we will show that the following vectors fulfill the conditions (up to a common factor in the components):

for the eigenvalues

\[ \begin{align}
\lambda_1 \:&=\: {1 \over 2} \left(\, \left(\alpha \,+\, \gamma \right) \,-\, \left[ \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2 \,\right]^{1/2} \, \right) \,, \\[8pt]
\lambda_2 \:&=\: {1 \over 2} \left(\, \left(\alpha \,+\, \gamma \right) \,+\, \left[ \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2 \,\right]^{1/2} \, \right) \,.
\end{align}
\]

As usual, the T at the formulas for the vectors symbolizes a transposition operation.

Note again that we have

\[
\lambda_1 \:\le\: \lambda_2 \,.
\]

This reflects our initial assumptions about the axes-lenghts of our ellipses. It means that the eigenvalue λ₁ and the respective eigenvector ξ₁ are associated with the longer half-axis of the ellipse! We had assumed that this half-axis is aligned with the x-coordinate axis.

Note also that the vector components given above are not normalized. This is important for performing numerical checks as Numpy and linear algebra programs would typically give you normalized eigenvectors with a length = 1. But you can easily compensate for this by working with

Proof for the eigenvalues and eigenvector components

We just prove that the eigenvector conditions are e.g. fulfilled for the components of the first eigenvector ξ₁ and λ₁ = λ_u.

\[ \begin{align}
\left(\alpha \,-\, \lambda_1 \right) * \xi_{1,x} \,+\, {1 \over 2} \beta * \xi_{1,y} \,&=\, 0 \\[8pt]
{1 \over 2} \beta * \xi_{1,x} \,+\, \left( \gamma \,-\, \lambda_1 \right) * \xi_{2,y} \,&=\, 0
\end{align}
\]

(The steps for the second eigenvector are completely analogous).

We start with the condition for the first component:

\[ \begin{align}
&\left( \alpha \,-\,
{1\over 2}\left[\,\left(\alpha \, + \, \gamma\right) \,-\, \left[ \beta^2 \,+\, \left( \alpha \,-\, \gamma \right)^2 \right]^{1/2} \right] \right) * \\[8pt]
& {1 \over \beta}\,
\left[\, \left(\alpha \,-\, \gamma\right) \,-\, \left[ \beta^2 \,+\, \left(\alpha \,-\, \gamma \right)^2 \right]^{1/2} \,\right] \,+\, {\beta \over 2 }
\,=\, 0
\end{align}
\]

\[ \begin{align}
& {1 \over 2 } \left[ \left(\alpha \,-\,\gamma\right) \,+\, \left[ \beta^2 \,+\, \left( \alpha \,-\, \gamma \right)^2 \right]^{1/2} \right] * \\[8pt]
& {1 \over \beta}\,
\left[\, \left(\alpha \,-\, \gamma\right) \,-\, \left[ \beta^2 \,+\, \left(\alpha \,-\, \gamma \right)^2 \right]^{1/2} \,\right] \,+\, {\beta \over 2 }
\,=\, 0 \,.
\end{align}
\]

\[
{1 \over 2 \, \beta} \left[ (\alpha \,-\,\gamma)^2 \,-\, \beta^2 \,-\, (\alpha \,-\,\gamma)^2 \right] \,+\, {\beta \over 2 } \,=\, 0
\]

The last relation is obviously true.

You can perform a similar calculation for the other eigenvector component:

\[ \begin{align}
{1 \over 2} \, \beta & {1 \over \beta}\,
\left[\, \left(\alpha \,-\, \gamma\right) \,-\, \left[ \beta^2 \,+\, \left(\alpha \,-\, \gamma \right)^2 \right]^{1/2} \,\right] \,+\, \\
&
\left( \gamma \, -\,
{1\over 2}\left[\,\left(\alpha \, + \, \gamma\right) \,-\, \left[ \beta^2 \,+\, \left( \alpha \,-\, \gamma \right)^2 \right]^{1/2} \right] \right) * 1 \,=\, 0
\end{align}
\]

Thus:

\[ \begin{align}
&{1 \over 2} \, \left(\alpha \,-\, \gamma\right) \,-\, {1 \over 2} \left[ \beta^2 \,+\, \left(\alpha \,-\, \gamma \right)^2 \right]^{1/2} \\
-\, &{1\over 2}\left(\alpha \,-\, \gamma\right) \,+\, {1\over 2}\left[ \beta^2 \,+\, \left( \alpha \,-\, \gamma \right)^2 \right]^{1/2} \,=\, 0
\end{align}
\]

True, again.

In a very similar exercise one can show that the scalar product of the eigenvectors is equal to zero:

\[ \begin{align}
& {1 \over \beta^2}\,
\left[\, \left(\alpha \,-\, \gamma\right)^2 \,-\, \beta^2 \,+\, \left(\alpha \,-\, \gamma \right)^2 /\,\right] \,+\, 1 \\[8pt]
& – {1 \over \beta^2} * \beta^2 \,*\,1 \,=\, 0 \,.
\end{align}
\]

I.e.,

\[
\pmb{\xi_1} \bullet \pmb{\xi_1} \,=\, \left( \xi_{1,x}, \, \xi_{1,y} \right) \circ \begin{pmatrix} \xi_{2,x} \\ \xi_{2,y} \end{pmatrix} \,= \, 0 \,.
\]

The eigenvectors are perpendicular to each other. Exactly, what we expect for the orientations of the principal axes of an ellipse against each other.

Rotation angle from coefficients of A_q

We still need a formula for the rotation angle(s). From linear algebra results related to an eigendecomposition we know that the orthogonal (rotation) matrices consist of columns of the normalized eigenvectors. With the components given in terms of our un-rotated CCS, in which we basically work. These vectors point along the principal axes of our ellipse. Thus the components of these eigenvectors define our aspired rotation angles of the ellipse’s principal axes against the x-axis of our CCS.

Let us prove this. By assuming

\[ \begin{align}
\cos (\phi_1) \,&=\, \xi_{1,x}^n \,, \\[8pt]
\sin (\phi_1) \,&=\, \xi_{1,y}^n
\end{align}
\]

and using

\[
\sin(2\phi_1) \,=\, 2\, \sin (\phi_1) \, \cos (\phi_1) \,,
\]

we get

\[ \begin{align}
\sin (2 \phi_1) \,=\,
2 * { \xi_{1,x} * \xi_{1,y} \over \left[\, \xi_{1,x}^2 \, + \, \xi_{1,y}^2 \,\right] } \,.
\end{align}
\]

Thus

\[ \begin{align}
\operatorname{sin}(2 \phi_1) \,&=\,
2 \,\, { {1 \over \large{\beta}} \left( (\alpha \,-\, \gamma) \,-\, \left[\, \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2\,\right]^{1/2} \right) \,*\, 1
\over
\left[\, \left( {1 \over \large{\beta}} \left( (\alpha \,-\, \gamma) \,-\, \left[\, \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2\,\right]^{1/2} \right) \right)^2
\,+\, 1^2 \right] } \\[8pt]
&=\, 2\,\, { {1 \over \large{\beta}} \left( t \,-\, z \right)
\over
{1 \over \large{\beta}^2 \phantom{\large{]}} } \left[\, \beta^2 \,+\, \left(\, t \,-\, z \,\right)^2 \right] } \,,
\end{align}
\]

with

\[ \begin{align}
t \,&=\, (\alpha \,-\, \gamma) \,, \\[8pt]
z \,&=\, \left[\, \beta^2 \,+\, \left(\gamma \,-\, \alpha \right)^2\,\right]^{1/2} \,.
\end{align}
\]

This looks very differently from the simple expression we got above. And a direct approach is cumbersome. The trick is to multiply nominator and denominator by a convenience factor

\[
\left( t \,+\, z \right),
\]

and exploit

\[ \begin{align}
\left( t \,-\, z \right) \, \left( t \,+\, z \right) \,&=\, t^2 \,-\, z^2 \,, \\[8pt]
\left( t \,-\, z \right) \, \left( t \,+\, z \right) \,&=\,\, – \,\beta^2
\end{align}
\]

to get

\[ \begin{align}
&2 * \beta \, { (t\,-\, z) * ( t\,+\, z) \over \left[ \beta^2 \, + \,( t \,-\, z )^2 \right] * (t \,+\, z) } \\[8pt]
&=\: 2 * \beta \, { – \beta^2 \over \beta^2 (t\,+\,z) \,-\, \beta^2 (t\,-\,z) } \\
&=\: – \, {\beta \over \left[\, \beta^2 \,+\, (\alpha \,-\, \gamma)^2 \,\right]^{1/2} } \,.
\end{align}
\]

This means that our 2nd approach gives us the result

\[
\operatorname{sin}\, (2 \phi_1) \:=\: \, – \, { \beta \over \left[\, \beta^2 \,+\, (\alpha \,-\, \gamma)^2 \,\right]^{1/2} }\,,
\]

which is of course identical to the result we got with our first solution approach. It is clear that the second axis has an inclination by φ +- π / 2:

\[
\phi_2\, =\, \phi_1 \,\pm\, \pi/2.
\]

In general the angles have a natural ambiguity of π. This makes life a bit harder when you get a matrix, use the formulas above and try to construct the matrix with correct orientation.

Addendum 07/05/2025: Getting the right orientation of the ellipses when constructing them via eigenvalues of a matrix A_q

The formulas for the rotation angle of the ellipse look simple. However, some pitfalls may await you, when you want to use the formula to construct your ellipses after having determined its half-axes from the eignevalues of a given matrix A_q. If you are not careful and forget to take into account our special assumptions on axis-length and orientation of the eigenvectors you may get the ellipses orientation wrong.

Most often this happens when working with contour ellipses of BivariateNormal Distributions [BVDs] and their central covariance matrices, which correspond to our A_q-matrix. In the case of BVDs the matrix coefficients have an association with the standard-deviations of the BVDs marginal distributions and this may lead to confusion.

The important points to take care of are the following:

The arcsin-function allows for multiple equivalent values – and we have to choose the right one. To do this you should analyze both of the key equations determining the angle Φ:
\[ \begin{align}
\gamma \,\, – \, \alpha \,&=\, \left( \, \lambda_2 \,-\, \lambda_1\,\right) \, \cos (2\,\phi) \,, \\[8pt]
\,-\, \beta \:&=\: \left( \, \lambda_2 \,-\, \lambda_1 \, \right)\, \sin (2\,\phi) \,.
\end{align}
\]

This will narrow down the interval in which Φ resides.
Throughout our text, we have used the assumption λ₁ = λ_u ≤ λ₂ = λ_d. A priori it is questionable whether this covers all necessary cases during the reconstruction of an ellipse properly.
Not all given matrices may fulfill the assumption about the λ-values. Instead, for certain matrices the eigenvalues λ₁ and λ₂ may switch their position in the central diagonal matrix of the eigendecomposition. This must be analyzed.
It can, however, be shown that cases with λ₁ < λ₂ can be mapped 1:1 onto certain cases with λ₁ > λ₂. A proper analysis is given in an extensive post in another blog.
The eigenvalue λ₁ and the respective eigenvector ξ₁ were in our considerations always associated with the longer half-axis of the ellipse! This has the consequence that the meaning of Φ₁ during reonstruction can change with respect to the real situation given by the matrix.

A thorough analysis of how the matrix elements determine the angle of the ellipse and respective recipes has meanwhile benn provided by myself in a related post in another blog. There you find also arguments why it is justified to focus on cases with (λ₂ – λ₁) ≥ 0.

Conclusion

In this post I have shown how one can derive essential properties of centered, but rotated ellipses from matrix-based representations. Such calculations become relevant when e.g. experimental or numerical data only deliver the coefficients of a quadratic form for the ellipse.

We have first established the relation of the coefficients of a matrix that defines an ellipse by a combined scaling and rotation operation with the coefficients of a matrix which defines an ellipse as a quadratic form of the components of position vectors. In addition we have shown how the coefficients of both matrices are related to quantities like the lengths of the principal axes of the ellipse and the inclination of these axes against the x-axis of the Cartesian coordinate system in which the ellipse is described via position vectors. So, if one of our two defining matrices is given we can numerically calculate the ellipse’s main properties.

Regarding the ellipses orientation we have to be careful and check which of its eigenvalues describes the longer axis.

In the next post of this mini-series

Properties of ellipses by matrix coefficients – II – coordinates of points with extremal y-values

we have a look at the x- and y-coordinates of points on an ellipse with extremal y-values. All in terms of the matrix coefficients we are now familiar with.

M	T	W	T	F	S	S
					1	2
3	4	5	6	7	8	9
10	11	12	13	14	15	16
17	18	19	20	21	22	23
24	25	26	27	28	29	30
31

Linux-Blog – Dr. Mönchmeyer / anracon

Notes about Linux, ML and some simple math …

Category Archives: Machine Learning

Fun with shear operations and SVD – V – matrices of sheared n-dimensional ellipsoids

Matrix describing a centered n-dimensional ellipsoid

Equation of the quadratic form for the sheared ellipsoid

Inclusion of a SVD eigendecomposition of M_S

An example for the case of a sheared ellipse

Conclusion

Fun with shear operations and SVD – IV – Shearing of ellipses

Already established results for shearing a circle

Objectives of this post: Shearing of a centered, rotated ellipse

Properties of ellipses by matrix coefficients – III – coordinates of points with extremal radii

Reduced matrix equation for an ellipse

Method 1 to determine the vectors to the principal axes’ end points

Equation for points with maximum radius values

Solution for x_e-values of the end-points of the principal axes

Plots

Conclusion

Properties of ellipses by matrix coefficients – II – coordinates of points with extremal y-values

Matrix equation for an ellipse

Points of the ellipse with extremal y_E – values

Solution in terms of the coefficients of an alternative matrix A_E

Conclusion

Properties of ellipses by matrix coefficients – I – Two defining matrices and eigenvalues

Centered ellipses and two related matrices

Matrix A_E of a centered and rotated ellipse: Scaling of a unit circle followed by a rotation

Quadratic forms – Case 1: Centered ellipse, principal axes aligned with CCS-axes

Quadratic forms – Case 2: General centered and rotated ellipse

Quadratic form: A matrix equation to define an ellipse

How to derive σ₁, σ₂ and φ from the coefficients of A_E or A_q in the general case?

Direct derivation of σ₁, σ₂ and φ from A_q by using trigonometric relations

Direct derivation of σ₁, σ₂ and φ from A_E by using trigonometric relations

Determination of the inclination angle φ

2nd way to a solution for σ₁, σ₂ and φ via eigendecomposition

Proof for the eigenvalues and eigenvector components

Rotation angle from coefficients of A_q

Addendum 07/05/2025: Getting the right orientation of the ellipses when constructing them via eigenvalues of a matrix A_q

Conclusion

Matrix describing a centered n-dimensional ellipsoid

Equation of the quadratic form for the sheared ellipsoid

Inclusion of a SVD eigendecomposition of MS

An example for the case of a sheared ellipse

Conclusion

Already established results for shearing a circle

Objectives of this post: Shearing of a centered, rotated ellipse

Reduced matrix equation for an ellipse

Method 1 to determine the vectors to the principal axes’ end points

Equation for points with maximum radius values

Solution for xe-values of the end-points of the principal axes

Plots

Conclusion

Matrix equation for an ellipse

Points of the ellipse with extremal yE – values

Solution in terms of the coefficients of an alternative matrix AE

Conclusion

Centered ellipses and two related matrices

Matrix AE of a centered and rotated ellipse: Scaling of a unit circle followed by a rotation

Quadratic forms – Case 1: Centered ellipse, principal axes aligned with CCS-axes

Quadratic forms – Case 2: General centered and rotated ellipse

Quadratic form: A matrix equation to define an ellipse

How to derive σ1, σ2 and φ from the coefficients of AE or Aq in the general case?

Direct derivation of σ1, σ2 and φ from Aq by using trigonometric relations

Direct derivation of σ1, σ2 and φ from AE by using trigonometric relations

Determination of the inclination angle φ

2nd way to a solution for σ1, σ2 and φ via eigendecomposition

Proof for the eigenvalues and eigenvector components

Rotation angle from coefficients of Aq

Addendum 07/05/2025: Getting the right orientation of the ellipses when constructing them via eigenvalues of a matrix Aq

Conclusion

Inclusion of a SVD eigendecomposition of M_S

Solution for x_e-values of the end-points of the principal axes

Points of the ellipse with extremal y_E – values

Solution in terms of the coefficients of an alternative matrix A_E

Matrix A_E of a centered and rotated ellipse: Scaling of a unit circle followed by a rotation

How to derive σ₁, σ₂ and φ from the coefficients of A_E or A_q in the general case?

Direct derivation of σ₁, σ₂ and φ from A_q by using trigonometric relations

Direct derivation of σ₁, σ₂ and φ from A_E by using trigonometric relations

2nd way to a solution for σ₁, σ₂ and φ via eigendecomposition

Rotation angle from coefficients of A_q

Addendum 07/05/2025: Getting the right orientation of the ellipses when constructing them via eigenvalues of a matrix A_q