← Home

SLAM and GTSAM

October, 2026

A few months ago, I found the time to simuatenously localize and map (SLAM). Though I have been close to SLAM impelementations for years, I have never built one from scratch by myself. Also as a depth camera lover--I know this is not a heathly feeling--I wanted to fully understand how come these SLAM systems are not relying on these lovely depth cameras. A sensor measuring 3D data must be super help for navigating 3D space, right?

My goal was to make a SLAM algorithm that run on the EuRoC MAV dataset with some reasonable level of accuracy. After some heavy bikeshedding with the data reading, and UI code, I saw the SIFT algorithm patent now got expired and felt some joy and justice there while writing some OpenCV. And then got to the stage for pose graph optimization, which was the main part I was waiting for. Pose graph optizmiation is, you build a graph with relative poses (translation + rotation) as edges, and then figuring out what is the best poses to give to each of the nodes of the graph, so the poses between the nodes matches the edges the best. It is some kind of a math problem and GTSAM is a solver for this math problem. And this problem space has yet been fully transformer-ized yet so, there is also that.

Honestly, I have expected to do more math, but GTSAM was actually really good at doing all the math. Getting the coordinate systems right was surprisingly hard--I don't know why all these modules need to have a different meaning to xyz? There were bunch of heuristics I have learned in matching feature points and picking the keyframes from incoming camera streams. And at the end, I got to a point the systems ATE error became around 3 cm for the dataset scenarios. This error is... around 10 times larger than the better systems, but as the work was increasingly starting to look like something I should get paid for doing, I have claimed this reasonable for now.

If you want to check out the code: https://github.com/hanseuljun/slam