My research uses language models as tools for studying human language and culture at scale. My work incorporates ideas from linguistic relativity, representation analysis, and geometric and topological methods.
These are not separate research interests so much as different perspectives on the same question: whether languages organize meaning and experience in systematically distinct ways, and whether those differences leave measurable traces in linguistic data and their latent representations.
(1) Linguistic Relativity / Linguistic Worldview
I take the soft, non-deterministic view of linguistic relativity (following the Lublin School of Ethnolinguistics) that languages make particular distinctions more habitual, salient, or obligatory to express. Languages differ in how they divide color space, lexicalize emotions, categorize objects, encode social relationships, and describe time, motion, agency, and evidence. The combined impact of these defining patterns forms culturally and historically distinct ways of construing experience — i.e., a linguistic worldview.
These phenomena have traditionally been studied through the humanistic, highly interpretive disciplines of ethnography, anthropology, and native-speaker intuition. Computational methods offer another scale of analysis. If culturally meaningful distinctions are systematic, they should leave traces in the collocational relationships between words and concepts.
(2) Representation Analysis / Interpretability
Language models form latent, internal representations of concepts from training on large collections of human language. These representations are not designed semantic systems, but rather emerge from the content of the training data and the environment of architecture and learning objectives.
Because models’ latent representations are shaped by the distributional patterns inherent in language, they become a new kind of empirical artifact for understanding how human concepts are organized. I am interested in how human-interpretable concepts are structured within latent space, and how stable or robust they are across architectures and languages.
I am most interested in approaches that involve:
(2.5) Topological Data Analysis / Geometric Data Analysis
Embedding spaces are high-dimensional, noisy, and difficult to compare. Because a semantic space is defined only by its content, its axes are arbitrary, and aligning two spaces to a shared frame is lossy or distortionary.
One way we can approximate is by evaluating a space by its geometric or topological invariants. My work mainly employs persistent homology, which tracks features that survive across continuous scaling.
Learn more about topological data analysis — coming soon!
(3) Language and Cognition
Any symbolic system (mathematics, computation, DNA, etc.) selects which distinctions are efficient to express and compresses those which are effortful. In doing so, it makes some features necessary and some nearly unexpressible. Language does not escape this constraint, and therefore its structure can tell us what the human mind optimizes for: patterns of attention and association.