NoteMathematicsRoboticsEvergreenupdated 2026.089 min read

Lie theory is the common language for carrying optimization over "curved spaces" — rotations, poses — back into ordinary vector calculus. This note walks from manifolds and tangent spaces to hat/vee and exp/log, assembles them into one optimization step on a manifold, and closes with a hand-written SO(3)SO(3) code sample.

A study note based on Aalok Patwardhan's A Visual Introduction to Lie Theory[1], with the concrete derivations and code filled in.

§1Prerequisites

The main text keeps returning to the following concepts, each written as an "English term = Chinese name = minimal definition" triple:

  • manifold = 流形 = a smoothly curved space that looks locally like flat Rn\mathbb{R}^n but is globally curved.[2]
  • Lie group = 李群 = an object that is both a manifold and a group whose operations (matrix multiplication and inversion) are smooth.[4]
  • tangent space = 切空间 = the flat tangent plane attached to the manifold at a given point.[2]
  • Lie algebra = 李代数 = the special tangent space of a Lie group at the identity element (written so(n)\mathfrak{so}(n) for rotation groups).[2]
  • skew-symmetric matrix = 反对称矩阵 = a matrix satisfying A⊤=−AA^\top=-A; elements of a rotation group's Lie algebra take exactly this form.[2]
  • exponential map / logarithm map = 指数映射 / 对数映射 = a pair of mutually inverse maps travelling between the tangent space and the manifold; for rotations they are the matrix exponential and matrix logarithm.[4]

§2Recap: optimization in Euclidean space

Start with the case where everything goes smoothly. Given a cost function f:Rn→Rf:\mathbb{R}^n\to\mathbb{R}, we want the x\mathbf{x} that minimizes it. The gradient descent recipe is:

  1. perturb: nudge x\mathbf{x} by a tiny amount;
  2. gradient: compute ∇f(x)\nabla f(\mathbf{x}), which points in the direction of steepest ascent;
  3. step: take a small step along −∇f-\nabla f.
xk+1=xk−α ∇f(xk)\mathbf{x}_{k+1} = \mathbf{x}_k - \alpha\,\nabla f(\mathbf{x}_k)

The crucial premise here is that x\mathbf{x} is a free vector. Each of its components can independently absorb a small increment δ\delta, and the result is still a valid input. Rn\mathbb{R}^n is "flat" and addition is unconstrained, so perturbing and stepping are entirely natural.

§3The dilemma: optimizing over a rotation

Take 2D rotations as an example; the rotation matrix is:

R(θ)=[cos⁡θ−sin⁡θsin⁡θcos⁡θ]R(\theta)=\begin{bmatrix}\cos\theta & -\sin\theta\\[2pt] \sin\theta & \cos\theta\end{bmatrix}

It has 4 entries, but they are not free — two constraints must hold simultaneously:

R⊤R=I(orthogonal: unit orthogonal columns),det⁡R=1(right-handed, no reflection)R^\top R = I \quad(\text{orthogonal: unit orthogonal columns}),\qquad \det R = 1\quad(\text{right-handed, no reflection})

Suppose we do what we would in R4\mathbb{R}^4 and add a small increment δ\delta to the top-left cos⁡θ\cos\theta:

[cos⁡θ+δ−sin⁡θsin⁡θcos⁡θ]\begin{bmatrix}\cos\theta+\delta & -\sin\theta\\ \sin\theta & \cos\theta\end{bmatrix}

This is exactly where ordinary optimization hits a wall: the valid rotations do not form a flat vector space but a curved, constrained subset. Doing calculus on it requires a different toolset.

§4Manifolds and Lie groups

The set of all valid 2D rotations is called the special orthogonal group SO(2)SO(2); the 3D version is SO(3)SO(3):

SO(n)={ R∈Rn×n ∣ R⊤R=I, det⁡R=1 }SO(n)=\{\,R\in\mathbb{R}^{n\times n}\ \mid\ R^\top R=I,\ \det R = 1\,\}

They are manifolds: smoothly curved spaces that locally look like flat Rn\mathbb{R}^n (much as a patch of the Earth's surface is well approximated by a flat map) but are globally curved. An object that is both a manifold and a group with smooth operations (matrix multiplication and inversion) is a Lie group[4].

The three spaces: manifold, tangent space, R^n

Figure 1. A point XX on a Lie group; its tangent space is a "plane" attached at that point, and that plane is in turn isomorphic to Rn\mathbb{R}^n. All the optimization happens on the far right.

§5Lie algebra and tangent space

At a point XX on the manifold, one can attach a flat tangent plane called the tangent space. The special tangent space at the identity element II is the Lie algebra of the Lie group, written so(n)\mathfrak{so}(n)[2].

For rotations, the elements of the Lie algebra turn out to be exactly the skew-symmetric matrices (反对称矩阵, A⊤=−AA^\top=-A):

so(2)\mathfrak{so}(2)

Only one free parameter:

θ∧=[0−θθ0]=θ[0−110]⏟G\boldsymbol{\theta}^\wedge=\begin{bmatrix}0 & -\theta\\ \theta & 0\end{bmatrix}=\theta\underbrace{\begin{bmatrix}0&-1\\1&0\end{bmatrix}}_{G}

so(3)\mathfrak{so}(3)

Three free parameters ω=(ω1,ω2,ω3)\boldsymbol{\omega}=(\omega_1,\omega_2,\omega_3):

ω∧=[0−ω3ω2ω30−ω1−ω2ω10]\boldsymbol{\omega}^\wedge=\begin{bmatrix}0 & -\omega_3 & \omega_2\\ \omega_3 & 0 & -\omega_1\\ -\omega_2 & \omega_1 & 0\end{bmatrix}

It doubles as the cross-product operator: ω∧v=ω×v\boldsymbol{\omega}^\wedge\mathbf{v}=\boldsymbol{\omega}\times\mathbf{v}.

Note that the number of independent components of a skew-symmetric matrix (1 for SO(2)SO(2), 3 for SO(3)SO(3)) is exactly the intrinsic dimension of the manifold. This is no accident — the dimension of the tangent space is the number of degrees of freedom. That hands us a flat, unconstrained space in which to do the math.

§6hat and vee: Rn\mathbb{R}^n as a proxy tangent space

The tangent space (those skew-symmetric matrices) is flat, but written as matrices it is still awkward to feed straight into an optimizer. Fortunately it is isomorphic to the ordinary vector space Rn\mathbb{R}^n — two mutually inverse operators ferry elements between them:

  • hat (⋅)∧: Rn→so(n)(\cdot)^\wedge:\ \mathbb{R}^n\to\mathfrak{so}(n): lifts a workspace vector into the Lie algebra;
  • vee (⋅)∨: so(n)→Rn(\cdot)^\vee:\ \mathfrak{so}(n)\to\mathbb{R}^n: flattens a Lie algebra element back into a workspace vector.
ω∈R3 → (⋅)∧  ω∧∈so(3) → (⋅)∨  ω∈R3,((ω∧))∨=ω\boldsymbol{\omega}\in\mathbb{R}^3 \ \xrightarrow{\ (\cdot)^\wedge\ }\ \boldsymbol{\omega}^\wedge\in\mathfrak{so}(3) \ \xrightarrow{\ (\cdot)^\vee\ }\ \boldsymbol{\omega}\in\mathbb{R}^3,\qquad \big((\boldsymbol{\omega}^\wedge)\big)^\vee=\boldsymbol{\omega}

So the object we actually hand to gradient descent is that plain 3-dimensional vector ω\boldsymbol{\omega} — it carries no constraints and can be perturbed however we like.

§7exp and log: connecting the curved and the flat

The last piece of the puzzle is the bridge that travels between the manifold and the tangent space:

  • exponential map exp⁡: m→M\exp:\ \mathfrak{m}\to\mathcal{M}: "wraps" an element of the tangent space back onto the curved manifold;
  • logarithm map log⁡: M→m\log:\ \mathcal{M}\to\mathfrak{m}: conversely, "unrolls" an element of the manifold onto the tangent space.
X=exp⁡(τ∧),τ∧=log⁡(X)X=\exp(\boldsymbol{\tau}^\wedge),\qquad \boldsymbol{\tau}^\wedge=\log(X)

For rotations, exp⁡/log⁡\exp/\log here are just the matrix exponential and matrix logarithm.

SO(2)SO(2): transparent at a glance

Substituting θ∧=θG\theta^\wedge=\theta G into the matrix exponential series and using G2=−IG^2=-I, the terms assemble themselves into sin⁡/cos⁡\sin/\cos:

exp⁡(θG)=I+θG+θ22!G2+⋯=cos⁡θ I+sin⁡θ G=[cos⁡θ−sin⁡θsin⁡θcos⁡θ]=R(θ)\exp(\theta G)=I+\theta G+\tfrac{\theta^2}{2!}G^2+\cdots =\cos\theta\,I+\sin\theta\,G =\begin{bmatrix}\cos\theta & -\sin\theta\\ \sin\theta & \cos\theta\end{bmatrix}=R(\theta)

Conversely log⁡R(θ)=θG\log R(\theta)=\theta G, i.e. θ=atan2⁡(R21,R11)\theta=\operatorname{atan2}(R_{21},R_{11}). The number θ\theta living in the tangent space is the rotation angle itself.

SO(3)SO(3): the Rodrigues formula

Let ω=θ u^\boldsymbol{\omega}=\theta\,\hat{\mathbf{u}}, where θ=∥ω∥\theta=\|\boldsymbol{\omega}\| is the rotation angle and u^\hat{\mathbf{u}} is the unit rotation axis. Using the so(3)\mathfrak{so}(3) identity (ω∧)3=−θ2 ω∧(\boldsymbol{\omega}^\wedge)^3=-\theta^2\,\boldsymbol{\omega}^\wedge to collapse the series gives Rodrigues' rotation formula[2][3]:

R=exp⁡(ω∧)=I+sin⁡θθ ω∧+1−cos⁡θθ2 (ω∧)2R=\exp(\boldsymbol{\omega}^\wedge)=I+\frac{\sin\theta}{\theta}\,\boldsymbol{\omega}^\wedge+\frac{1-\cos\theta}{\theta^2}\,(\boldsymbol{\omega}^\wedge)^2

The inverse (log map):

θ=arccos⁡ ⁣(tr⁡(R)−12),ω∧=log⁡(R)=θ2sin⁡θ (R−R⊤)\theta=\arccos\!\Big(\frac{\operatorname{tr}(R)-1}{2}\Big),\qquad \boldsymbol{\omega}^\wedge=\log(R)=\frac{\theta}{2\sin\theta}\,\big(R-R^\top\big)

§8Putting it together: one optimization step on a manifold

Now chain the three spaces into a closed loop. Let the cost function f(X)f(X) be defined on the manifold (X∈SO(3)X\in SO(3)). The key trick is to use a right perturbation to parameterize XX by a local, unconstrained small vector τ∈R3\boldsymbol{\tau}\in\mathbb{R}^3:

X(τ)=X exp⁡(τ∧)  ≡  X⊞τX(\boldsymbol{\tau}) = X\,\exp(\boldsymbol{\tau}^\wedge)\;\equiv\;X\boxplus\boldsymbol{\tau}

(The ⊞\boxplus notation follows the convention of micro Lie theory[2].)

Take the gradient of ff with respect to τ\boldsymbol{\tau} at τ=0\boldsymbol{\tau}=\mathbf{0} (this step happens entirely inside flat R3\mathbb{R}^3, with the ordinary chain rule), obtain the gradient g∈R3\mathbf{g}\in\mathbb{R}^3, and then:

This is the diagram of the whole note, made concrete: log to flatten → optimize in Rn\mathbb{R}^n → exp to wrap back. The constraints are guaranteed automatically by exp⁡\exp, and all the optimizer ever sees is one free little vector.

§9Code sample: hand-writing hat / vee / exp / log for SO(3)SO(3)

NumPy implementation

python
import numpy as np

def hat(w):                      # R^3 -> so(3)
    wx, wy, wz = w
    return np.array([[0, -wz,  wy],
                     [wz,  0, -wx],
                     [-wy, wx,  0]])

def vee(W):                      # so(3) -> R^3
    return np.array([W[2, 1], W[0, 2], W[1, 0]])

def exp_so3(w):                  # Rodrigues: R^3 -> SO(3)
    theta = np.linalg.norm(w)
    W = hat(w)
    if theta < 1e-8:             # θ→0: fall back to Taylor
        return np.eye(3) + W
    a = np.sin(theta) / theta
    b = (1 - np.cos(theta)) / theta**2
    return np.eye(3) + a * W + b * (W @ W)

def log_so3(R):                  # SO(3) -> R^3
    theta = np.arccos(np.clip((np.trace(R) - 1) / 2, -1.0, 1.0))
    if theta < 1e-8:
        return vee(R - np.eye(3))
    return theta / (2 * np.sin(theta)) * vee(R - R.T)

Self-check

python
w = np.array([0.3, -0.7, 1.1])
R = exp_so3(w)
assert np.allclose(R.T @ R, np.eye(3))        # still orthogonal
assert np.allclose(np.linalg.det(R), 1.0)     # det = 1
assert np.allclose(log_so3(R), w)             # log ∘ exp = id

Off-the-shelf libraries

In a real project, don't reinvent the wheel — reach for a mature implementation:

  • Sophus (C++, SO3/SE3, Jacobians included)
  • manif (C++/Python, the reference implementation of micro Lie theory[2])
  • GTSAM / Ceres (wrap manifold optimization as factor graphs / a Manifold type)

§10Why all this matters

Lie theory is the lingua franca of modern robotic state estimation[2][3]. Almost anywhere rotations or poses have to be optimized or integrated, it is at work:

SettingWhere Lie theory enters
SLAM / bundle adjustmentCamera poses live in SE(3)SE(3); Gauss–Newton runs in the tangent space
pose graph optimizationNodes are poses, edges are relative constraints; residuals and Jacobians are computed in the Lie algebra
IMU preintegrationA gyroscope measures angular velocity; integration happens on SO(3)SO(3) rather than by Euclidean accumulation
state estimation / EKFUncertainty is modelled as a Gaussian in the tangent space (error-state Kalman filter)

§11References

  1. Aalok Patwardhan, A Visual Introduction to Lie Theory, aalok.uk interactive tutorial — the original source of this note.
  2. J. Solà, J. Deray, D. Atchuthan, A micro Lie theory for state estimation in robotics, arXiv:1812.01537 (2018) — the authoritative short treatment from an engineering perspective, with manif as its companion library.
  3. T. D. Barfoot, State Estimation for Robotics, Cambridge University Press (2017) — a systematic textbook; Chapter 7 covers SO(3)/SE(3)SO(3)/SE(3).
  4. B. C. Hall, Lie Groups, Lie Algebras, and Representations: An Elementary Introduction, 2nd ed., Springer, Graduate Texts in Mathematics 222 (2015).
Titles, sections and body text, in this language.
    ↑↓ · Enter · Escastro-inkstone