Decide the Role Before You Read the Track Record
Seven gates for adding a QIS convexity sleeve: naming the mandate, defining drawdown states, building peer-based expectations and testing stress liquidity.
9 min read | Aug 24, 2026
Most reviews of a systematic convexity sleeve start with the strategies and work outward. That order is wrong. The constraints that decide the outcome nearly all follow from a question about role, and that question usually gets answered last, if it gets answered at all.
What follows is a sequence of seven gates, and the order is the argument. Fail an early gate and the later analysis is worthless, not merely incomplete. A portfolio optimised before the mandate is settled has been optimised against the wrong objective, and no amount of later rigour recovers that.
Name the mandate
Three different jobs are routinely given the same name.
|
Mandate |
What it buys |
What it demands of the design |
|
Drawdown participation |
Payoff in falling equity months |
Participation in the states that actually occur; standalone risk-adjusted return is not the objective |
|
Diversification |
Return not already owned elsewhere in the book |
Spanning tests against current holdings; low pairwise correlation is not sufficient |
|
Funding liquidity |
Assets saleable at the moment capital is called |
Settlement terms, issuer exposure, behaviour of the bid in stress |
The three mandates imply different portfolios. A sleeve built for the first will not necessarily satisfy the third.
For a book with a substantial illiquid allocation, the third mandate is frequently the binding one, and it is the one least often written down. Capital calls are contractual and countercyclical: general partners call harder when assets are cheap. Distributions are discretionary and pro-cyclical, and thin out in the same window. The relationship is directional and not mechanical, and 2022 is the obvious caution, since calls slowed and distributions collapsed together. But the asymmetry that matters survives that episode: a call schedule is an obligation and a distribution schedule is not.
Net illiquid outflow tends to peak precisely when liquid funding capacity troughs.
None of that requires a correlation estimate. It is balance-sheet mechanics, which is why the mandate question comes first.
Define the drawdown states before opening a track record
Partition equity months into calm, shallow and deep, and fix the thresholds in advance. Conditional statistics chosen after the returns have been seen are decoration.
Deep months are rare. Shallow ones are not. Cumulative loss is frequency multiplied by severity, and for most portfolios over a realistic holding period the drag from repeated shallow episodes is the larger of the two terms. Deep-crash protection is nonetheless the more heavily marketed product and the more expensive one, and overshooting a deep-tail floor is paid for continuously in carry, in exchange for a state that may arrive once a decade. So the band the sleeve is bought for is a design choice, and it should be settled before anyone opens a track record.
Build an expectation, not a haircut
The standard complaint about backtests is that they are too good. That describes the problem without diagnosing it. A simulated record is a single observation drawn from a selected sample, and the selection is not random: strategies tend to be brought to market after a favourable run for the return driver they depend on. Much of what follows in live trading is regression to the mean. And regression to the mean says something about a distribution, which means you need to have one.
To shrink an observed result toward a mean, you have to know what the mean is. Without a peer group there is nothing to regress toward, and the simulated record gets judged against zero or against its own history. Neither tells you how a strategy of this type, launched at this point in the cycle, has tended to behave afterwards. Much of the disappointment attributed to backtesting is really a missing reference distribution.
Building one is exacting. Four requirements matter.
-
Group by return driver, not by label or by provider. Two strategies with the same name frequently monetise different things.
-
Live data only, with closed and discontinued strategies retained. A distribution built from what survived measures survival.
-
Cohort by launch date, so a candidate sits against peers that started into the same conditions and not against the whole history of the category.
-
Comparable risk. A distribution spanning different volatility levels compares nothing.
Cohorts thin out quickly, which is the fair objection to the third one. The answer is to widen the window until the sample supports an estimate, then report the width you used. That is better than dropping the conditioning and comparing against everything. A coarse conditional figure with a stated sample size beats a precise unconditional one.
With that in place the expectation becomes conditional instead of arithmetic. Two candidates presenting identical simulated records deserve different expectations if one launched into a favourable regime for its driver and the other did not. How large the adjustment should be depends on how extreme the result looks against peers and on where in the cycle it was produced. It is not a constant you subtract from everything.
One asymmetry carries forward. Simulated risk statistics travel much better than simulated returns: volatility and tail measures from a backtest are broadly usable, the return figure is not. So the same file is a reasonable input to sizing and a poor input to expectation, and it should be read twice with different degrees of trust. Assembling the reference distribution is the expensive part of all this, and it cannot be borrowed from the provider, whose peer group is its own shelf.
Fix the constraint, then optimise the underlying
Leverage is an input to the design and not an output of it. Fix the gross exposure of the wrapper first, then optimise the components against the payoff the portfolio actually needs.
The alternative is more common: optimise components on standalone risk-adjusted return, then lever the result to the exposure required. It produces a different book, because the two procedures reward different things. An unlevered optimisation favours components with high return per unit of risk, which tends to select carry. Fixing the exposure first and optimising for participation in the states defined at gate 2 favours components that pay in those states, which tends to select convexity, and convexity looks expensive on a standalone basis. Both books can be presented at the same gross exposure and the same volatility target, and they will behave differently in the states the sleeve was bought for.
Judge marginal contribution, never standalone merit
The test for a candidate is what it does to the book already owned: correlation to it, uniqueness measured as one minus R-squared against it, and the movement in volatility, conditional value at risk and the conditional figures from gate 2 at a given weight. A strategy with an unremarkable standalone record can be the right position. A strong one can be redundant. Judging strategies on their own merits is the more common approach and the wrong one.
Rotation is why this is a recurring test. Top-quartile persistence in this asset class falls to 26% after a year, so the sleeve being underwritten is not the sleeve that will be held. Turnover here is a structural requirement, not a house style, and every rebalance re-poses the marginal contribution question against a book that has itself changed. A programme that runs the test once at entry is measuring a portfolio it no longer owns.
Test whether it can be monetised in the window
A hedge that cannot be sold in the week it is needed is a hedge on the factsheet only.
This is where liquid convexity earns its keep, and where wrapper detail decides whether it does. Four things belong on paper before the position is sized.
-
Redemption terms and any gating provisions, tested against the horizon over which you would actually sell the sleeve and not the horizon in the marketing document.
-
Issuer credit exposure, which correlates with the conditions under which the position would be monetised. Counterparty count should be chosen, not inherited.
-
Behaviour of the bid in stress. The calm figure is the one that appears in the term sheet.
-
Execution format, which turns out to be worth real money.
Count the events the case rests on
Conditional statistics drawn from a live sample since 2018 typically rest on a handful of deep-drawdown months, most of them belonging to a single episode. Small samples, wide error bars. State how many observations support each conditional figure, and treat a deep-tail number resting on one crash as indicative and not as a probability-weighted expectation.
Shallow-band statistics rest on considerably more observations, which is a further argument for building around them. Gates 2 and 7 meet here: the band that is cheapest to hedge is also the band the data can speak to.
The disqualifying answers
A framework is only useful if it can return a negative. The sleeve should not be added when:
-
the mandate cannot be stated in a single sentence, because the portfolio will then be built against the wrong objective;
-
the case rests on simulated history with no peer distribution to set it against, which leaves the expectation unanchored;
-
funding liquidity is the real mandate but the wrapper carries terms that fail in the state where it would be used;
-
the position is being sized on standalone metrics, which means gate 5 was skipped rather than passed;
-
the deep-tail figure is the headline and rests on a single episode.
What the sequence produces
|
Gate |
The artefact it leaves behind |
|
1 Mandate |
One sentence stating what the sleeve is for |
|
2 Drawdown states |
Thresholds fixed in advance of any track record |
|
3 Expectation |
A peer distribution the simulated record is shrunk toward |
|
4 Constraint |
Gross exposure set as an input to the optimisation |
|
5 Marginal contribution |
Uniqueness and risk movement against the existing book, re-tested at each rebalance |
|
6 Monetisation |
Redemption terms, issuer exposure, stress bid, execution format |
|
7 Observation count |
Number of events supporting each conditional figure |
Every one of those artefacts is written down, testable against history, and capable of being shown to be wrong. None of them requires a view on anybody's skill. As with multi-manager platforms, the parts of this decision that are hardest to observe are not the parts doing most of the work.
Resonanz insights in your inbox...
Get the research behind strategies most professional allocators trust, but almost no-one explains.