Publications
Block construction is ubiquitous in early development, yet is surprisingly complex, involving stepby-step sequenced actions to create specific structures. Here, we use novel analytic methods to characterize these action sequences in detail, including which individual parts of the structure ('states') are built and how these structures are combined, creating a fully specified build path towards the final structure. We find that, like adults tested in a previous study, 4- to 8-year-olds build by creating a small subset of possible individual states and full build paths, and that they prioritize building layer-by-layer. The individual states and build paths that children produce are strikingly similar to those of adults, resulting in structures that are more stable than other possible (but not attested) states and paths. Our approach serves as a lens into the cognitive processes underlying block building and suggests that children's building is guided by significant cognitive constraints consistent with computational thinking.
Spatial construction-the activity of creating novel spatial arrangements or copying existing ones-is a hallmark of human spatial cognition. Spatial construction abilities predict math and other academic outcomes and are regularly used in IQ testing, but we know little about the cognitive processes that underlie them. In part, this lack of understanding is due to both the complex nature of construction tasks and the tendency to limit measurement to the overall accuracy of the end goal. Using an automated recording and coding system, we examined in detail adults' performance on a block copying task, specifying their step-by-step actions, culminating in all steps in the full construction of the build-path. The results revealed the consistent use of a structured plan that unfolded in an organized way, layer by layer (bottom to top). We also observed that complete layers served as convergence points, where the most agreement among participants occurred, whereas the specific steps taken to achieve each of those layers diverged, or varied, both across and even within individuals. This pattern of convergence and divergence suggests that the layers themselves were serving as the common subgoals across both inter and intraindividual builds of the same model, reflecting cognitive chunking. This structured use of layers as subgoals was functionally related to better performance among builders. Our findings offer a foundation for further exploration that may yield insights into the development and training of block-construction as well as other complex cognitive-motor skills. In addition, this work offers proof-of-concept for systematic investigation into a wide range of complex action-based cognitive tasks.
In this letter we address the task of recognizing assembly actions as a structure (e.g. a piece of furniture or a toy block tower) is built up from a set of primitive objects. Recognizing the full range of assembly actions requires perception at a level of spatial detail that has not been attempted in the action recognition literature to date. We extend the fine-grained activity recognition setting to address the task of assembly action recognition in its full generality by unifying assembly actions and kinematic structures within a single framework. We use this framework to develop a general method for recognizing assembly actions from observation sequences, along with observation features that take advantage of a spatial assembly's special structure. Finally, we evaluate our method empirically on two application-driven data sources: 1) An IKEA furniture-assembly dataset, and 2) A block-building dataset. On the first, our system recognizes assembly actions with an average framewise accuracy of 70% and an average normalized edit distance of 10%. On the second, which requires fine-grained geometric reasoning to distinguish between assemblies, our system attains an average normalized edit distance of 23%-a relative improvement of 69% over prior work.
Many applications of computer vision require robust systems that can parse complex structures as they evolve in time. Using a block construction task as a case study, we illustrate the main components involved in building such systems. We evaluate performance at three increasingly-detailed levels of spatial granularity on two multimodal (RGBD + IMU) datasets. On the first, designed to match the assumptions of the model, we report better than 90% accuracy at the finest level of granularity. On the second, designed to test the robustness of our model under adverse, real-world conditions, we report 67% accuracy and 91% precision at the mid-level of granularity. We show that this seemingly simple process presents many opportunities to expand the frontiers of computer vision and action recognition.


