I thought that too, but then they have some ambiguous sentences like 'In this paper we target high detail reconstruction from a single video captured in the wild, i.e., under uncontrolled imaging conditions,' where it seems as if they treat that as their input library. Other parts of the paper support your interpretation. It's annoyingly vague.
...but still, outstanding work even if it requires manual calibration like curation of the input set.