Joe Blog

How to Benchmark Private Equity Fund Performance: A Step-by-Step Guide

Written by Peter Harris, Investment Research Associate | Aug 11, 2026, 11:45:00 AM

Data sourced from Joe, the private fund performance platform powered by Dakota. Learn More | Request Access

Benchmarking a private equity fund correctly has become one of the most valuable capabilities in finance, and most GPs and LPs are still doing it wrong. A single IRR figure means almost nothing on its own. Without the right peer group, the right vintage-year context, and the right metric mix, a 14% Net IRR could be a strong result or a mediocre one, and there is no way to know which.

In this article, we're walking through the methodology step by step: which metrics to use, how to build a peer group that actually reflects a fund's competitive set, and the mistakes that make most benchmarking exercises misleading.

The Metrics That Actually Matter

Before comparing a fund to anything, get the metric selection right. Private equity performance is reported across several standard measures, and each tells a different part of the story.

  • Net IRR (Internal Rate of Return): The primary return metric, and the one most often quoted in isolation. It is money-weighted, which makes it sensitive to the size and timing of cash flows, not just the total amount returned.
  • TVPI (Total Value to Paid-In): Realized plus unrealized value as a multiple of capital invested. This captures total value creation regardless of timing.
  • DPI (Distributions to Paid-In): Realized returns only, cash actually returned to LPs. DPI is the metric least prone to manipulation because it measures money that has actually left the fund and landed in an LP's account.
  • RVPI (Residual Value to Paid-In): Unrealized NAV as a multiple of capital invested. High RVPI on an older fund can be a yellow flag: it may mean the GP is holding assets rather than realizing gains.
  • PME (Public Market Equivalent): Return versus an equivalent public market investment, calculated using methods like Kaplan-Schoar or Direct Alpha. PME answers the question LPs actually care about: did this fund beat what I could have earned in public markets over the same period.

No single metric tells the full story. A fund with a strong TVPI but weak DPI has created value on paper that has not yet been realized. A fund with strong Net IRR early in its life may simply be riding the mathematical effect of a fast initial distribution, not superior underwriting.

Why Vintage Year Changes Everything

Comparing a 2018 vintage fund's IRR to a 2023 vintage fund's IRR without adjustment is close to meaningless. Every private equity fund follows a J-curve: fees and early losses depress returns in the first several years, before realized gains catch up and eventually overtake them. A fund three years into its life has not had time to show what it can do, and its interim IRR can look artificially low or, in some cases, artificially high depending on the timing of one early exit.

This is why benchmark data is built and reported by vintage-year cohort rather than as one blended universe. A fund's performance only becomes meaningful once compared against other funds that started investing capital around the same time, since those funds faced the same entry pricing, the same exit environment, and the same macro conditions along the way.

Cohorts also need enough funds in them to mean something. A benchmark built from a handful of funds in a single vintage year can be swung heavily by one outlier, so the reliability of a benchmark depends as much on the size of its peer group as on the calculation behind it.

Building the Right Peer Group

Vintage year is the starting point, not the finish line. A benchmark that only controls for vintage still lumps together funds with fundamentally different return profiles. A software-focused middle market buyout fund and an industrials-focused middle market buyout fund from the same vintage face different growth dynamics, different multiple expansion environments, and different exit cycles.

A properly constructed peer group layers several filters at once:

  1. Asset class: Private equity, venture capital, private credit, private real estate, or real assets and infrastructure, since these have entirely different risk and return characteristics.
  2. Strategy or sub-strategy: Small market buyout, middle market buyout, large market buyout, or growth equity within private equity, since capital structure and hold periods vary meaningfully across these.
  3. Vintage year: Grouped into years or ranges (for example, 2019 through 2021) rather than a single year, to build a large enough sample while keeping the market conditions comparable.
  4. Geography: North America, Europe, Asia-Pacific, or global, since exit markets and valuation multiples differ by region.
  5. Portfolio company sector: Software, industrials, healthcare, energy, consumer, and other sectors that reflect what the fund's underlying companies actually do.

That fifth filter, sector of the underlying portfolio companies, is the one most benchmarking tools skip entirely. Most benchmark products stop at asset class, strategy, and vintage year. Filtering further by the sector composition of a fund's actual holdings is what separates a benchmark that reflects a fund's true competitive set from one that simply reflects its category label.

Building this kind of peer group by hand, filter by filter, is exactly what Joe automates. Joe, powered by Dakota, lets you stack asset class, strategy, vintage year, geography, and portfolio company sector in one query, and see quartile rankings update instantly. Request access to build your own peer group.

A Worked Example

Joe's benchmarking data illustrates how this comes together in practice. The platform tracks private fund performance by vintage-year cohort, reporting Net IRR, TVPI, and DPI at the top decile, top quartile, median, bottom quartile, and bottom decile for each cohort, built quarterly from fund-level performance records.

Two things stand out in cohort data structured this way. First, dispersion within a single vintage year is wide: the gap between top-decile and bottom-decile performance in a given vintage is often large enough that a fund's raw IRR alone says very little about how it actually performed relative to peers. Second, dispersion is not stable across vintages: cohorts from certain years show tighter spreads between top and bottom performers than others, reflecting how much market conditions during a fund's investment period shape outcomes across an entire cohort, not just individual funds within it.

Younger vintages (funds still in their first year or two) are typically excluded from quartile reporting altogether, marked not meaningful, since there has not been enough time for realized performance to emerge from the J-curve.

Common Mistakes That Undermine a Benchmark

  • Blending vintages. Combining a 2017 fund with a 2022 fund into one number erases the market-timing effects that made each fund's environment different.
  • Using a single metric. A benchmark built on Net IRR alone misses whether returns have actually been realized (DPI) or are still sitting on paper (RVPI).
  • Ignoring peer group size. A quartile ranking built from five comparable funds is far less reliable than one built from fifty, even if the filter criteria are identical.
  • Skipping sector-level filtering. Stopping at "middle market buyout" without accounting for what those funds actually invest in produces a peer group that includes fundamentally different return profiles under one label.
  • Presenting gross returns as net. Fee and carry treatment must be consistent across every fund in a peer group, or the comparison is not measuring the same thing.

What This Means for LP Due Diligence and GP Reporting

For LPs, a properly constructed benchmark is a diligence tool: it separates GPs who are genuinely top-quartile in their true peer group from those who only look strong because they were compared against a peer group that was too broad or too generic. For GPs, the same benchmark is a reporting tool: showing a fund's performance against a peer group defined by vintage, strategy, geography, and sector gives LPs a comparison they can actually examine and trust, rather than a black-box industry average.

 

Both sides benefit from the same underlying discipline: build the peer group first, then read the numbers. A benchmark's credibility comes from the transparency of its construction as much as from the accuracy of its math.

Joe's benchmarking data covers 18,000+ private funds across seven asset classes — private equity, venture capital, private credit, private real estate, infrastructure and real assets, hedge funds, and evergreen and interval funds — filterable by strategy, vintage, geography, and portfolio company sector.

Request access to Joe to see how a fund benchmarks against a peer group built from its actual competitive set.