No history yet

Introduction to Factor Models

Finding the Hidden Drivers

Why do the prices of different tech stocks often move together? Or why does a student who excels in algebra also tend to do well in geometry and calculus? At first glance, these might seem like separate events. But often, there are hidden forces, or factors, influencing them all at once.

Factor models are statistical tools that help us uncover these hidden drivers. They work by taking a large set of data that we can see and measure—like stock prices or test scores—and boiling it down to a few essential, unobservable characteristics. The goal is to simplify complexity and understand the underlying structure of our data.

Instead of tracking dozens of variables, a factor model might reveal that just two or three hidden factors are responsible for most of the patterns we see.

Observed vs. Latent

The core of any factor model lies in the distinction between two types of variables.

Observed variables are the things we can directly collect and measure. Think of a company's quarterly earnings, the daily temperature, or a person's answers on a survey. They are the raw data points we start with.

Latent factors, on the other hand, are the unobservable concepts we believe are causing those observations. We can't measure them directly with a ruler or a scale. Instead, we infer their existence from the patterns in the observed data.

latent

adjective

Existing but not yet developed or manifest; hidden or concealed.

For example, we can't directly measure "customer satisfaction." But we can measure related things like repeat purchases, survey ratings, and online reviews. In a factor model, these observed variables would help us estimate the strength of the underlying latent factor: customer satisfaction.

How It All Fits Together

The basic idea of a factor model is that each of our observed variables can be explained as a combination of these shared latent factors, plus a little bit of unique randomness or error.

The relationship between a factor and an observed variable is measured by something called a factor loading. A high loading means the factor has a strong influence on that variable. A low loading means the connection is weak.

Think of it like a recipe. The observed variable is the final dish. The latent factors are the main ingredients, and the factor loadings tell you how much of each ingredient to add. The error term is like a slight variation in cooking time that makes each dish unique.

Applications Across Fields

Factor models are incredibly versatile and show up in many different areas.

  • Finance: Investors use them to understand what drives stock returns. Instead of analyzing thousands of individual stocks, they might look at underlying factors like the overall market movement, company size, or industry trends.
  • Psychology: Researchers use factor analysis to develop personality tests. Your answers to dozens of questions (observed variables) might be used to measure a few underlying personality traits (latent factors) like extroversion or conscientiousness.
  • Economics: Economists might use data on employment, manufacturing output, and consumer spending (observed variables) to construct an index of "economic health" (a latent factor).
Lesson image

In each case, the goal is the same: to find a simpler, more meaningful story within a complex set of data.

Ready to check your understanding? This quiz will cover the core ideas we've just discussed.

Quiz Questions 1/5

What is the primary purpose of a factor model?

Quiz Questions 2/5

A market researcher wants to understand the latent factor of "brand loyalty." Which of the following would be an observed variable used to measure it?

Understanding these basic concepts is the first step. Next, we'll explore different types of factor models and how they are built.