Traders actively using backtesting often face a problem: after running dozens of strategies on historical data, there is no unified way to objectively compare them. Each strategy outputs its own set of metrics — one shows Sharpe ratio, another only maximum drawdown. Choosing the best becomes guesswork. We developed a system that standardizes all results: converts metrics to annual values, calculates additional indicators (Sortino, Calmar, profit factor), and outputs a summary table with rankings by each criterion. Thanks to a single StrategyResult class, all data is stored in one place, and the comparator automatically builds a ranking and analyzes strategy correlation.
Standardization is necessary because metrics calculated on different timeframes are incomparable. Sharpe on hourly candles and on daily candles gives different values. Our tool annualizes all metrics, as required by academic standards. This allows comparing strategies with different trade frequencies and test period lengths.
Why Metric Standardization is Critical for Backtesting?
Without a unified format, it's difficult to determine which strategy is truly better. One strategy's Sharpe ratio may be calculated on daily data, another on hourly data. Our system converts all metrics to annual values using a single StrategyResult class.
from dataclasses import dataclass
import pandas as pd
import numpy as np
@dataclass
class StrategyResult:
name: str
params: dict
equity_curve: pd.Series
trades: pd.DataFrame
# Computed metrics
sharpe_ratio: float
sortino_ratio: float
calmar_ratio: float
annual_return_pct: float
max_drawdown_pct: float
win_rate: float
profit_factor: float
total_trades: int
avg_trade_duration_hours: float
total_commission_pct: float
The class stores all computed metrics in one place. This makes it easy to pass results to the comparator.
How We Evaluate Risk and Return?
Each metric answers its own question: Sharpe shows excess return per unit of risk, Sortino considers only downside volatility, Calmar the ratio of return to maximum drawdown. We calculate them automatically from the equity curve.
| Metric | What It Shows | Formula (annualized) |
|---|---|---|
| Sharpe ratio | Return per risk | (mean_return - risk_free) / std_return * sqrt(252) |
| Sortino ratio | Return per downside risk | (mean_return - risk_free) / downside_std * sqrt(252) |
| Calmar ratio | Return to drawdown | annual_return / max_drawdown |
| Profit factor | Win/loss ratio | gross_profit / gross_loss |
| Win rate | Percentage of winning trades | wins / total_trades |
The table helps quickly understand which strategy is better for a given criterion.
Comparing Strategies with a Benchmark
We add a Buy & Hold benchmark to assess whether the strategy outperforms passive investing. The StrategyComparator ranks strategies by each metric and outputs an overall ranking.
class StrategyComparator:
def __init__(self, backtester, benchmark_data: pd.Series = None):
self.backtester = backtester
self.benchmark = benchmark_data # Buy & Hold BTC for comparison
def compare(self, strategies: list[dict], data: pd.DataFrame) -> ComparisonReport:
results = []
for strategy_config in strategies:
result = self.backtester.run(
strategy_class=strategy_config['class'],
params=strategy_config['params'],
data=data,
name=strategy_config['name'],
)
results.append(result)
if self.benchmark is not None:
bh_return = (self.benchmark.iloc[-1] / self.benchmark.iloc[0] - 1)
results.append(self._create_buyhold_result(self.benchmark))
return self.build_report(results)
def build_report(self, results: list[StrategyResult]) -> ComparisonReport:
comparison_df = pd.DataFrame([{
'Strategy': r.name,
'Annual Return %': round(r.annual_return_pct, 2),
'Sharpe Ratio': round(r.sharpe_ratio, 3),
'Sortino Ratio': round(r.sortino_ratio, 3),
'Max Drawdown %': round(r.max_drawdown_pct, 2),
'Calmar Ratio': round(r.calmar_ratio, 3),
'Win Rate %': round(r.win_rate * 100, 1),
'Profit Factor': round(r.profit_factor, 2),
'Total Trades': r.total_trades,
'Avg Trade Hours': round(r.avg_trade_duration_hours, 1),
'Commission Drag %': round(r.total_commission_pct, 2),
} for r in results])
rankings = self._compute_rankings(comparison_df)
correlations = self._compute_correlations(results)
return ComparisonReport(
summary=comparison_df,
rankings=rankings,
correlations=correlations,
strategies=results,
)
def _compute_rankings(self, df: pd.DataFrame) -> pd.DataFrame:
rankings = pd.DataFrame({'Strategy': df['Strategy']})
for metric, ascending in [
('Annual Return %', False),
('Sharpe Ratio', False),
('Max Drawdown %', True),
('Profit Factor', False),
('Win Rate %', False),
]:
if metric in df.columns:
rankings[f'Rank: {metric}'] = df[metric].rank(ascending=ascending).astype(int)
rank_cols = [c for c in rankings.columns if c.startswith('Rank:')]
rankings['Overall Rank'] = rankings[rank_cols].mean(axis=1).rank().astype(int)
return rankings.sort_values('Overall Rank')
def _compute_correlations(self, results: list[StrategyResult]) -> pd.DataFrame:
returns_dict = {
r.name: r.equity_curve.pct_change().dropna()
for r in results
}
returns_df = pd.DataFrame(returns_dict).dropna()
return returns_df.corr()
Ranking quickly identifies the leader by a composite of metrics, and the correlation matrix evaluates the degree of diversification.
How Visualization Helps Interpret Results?
Equity curve graphs, risk-return scatter plots, and drawdown charts provide a clear picture. Example code with Plotly:
def plot_comparison(report: ComparisonReport):
import plotly.graph_objects as go
from plotly.subplots import make_subplots
fig = make_subplots(
rows=2, cols=2,
subplot_titles=[
'Equity Curves',
'Risk-Return Scatter',
'Monthly Returns Distribution',
'Drawdown Comparison',
]
)
colors = ['#00C853', '#2196F3', '#FF9800', '#E91E63', '#9C27B0']
for i, strat in enumerate(report.strategies):
color = colors[i % len(colors)]
fig.add_trace(go.Scatter(
x=strat.equity_curve.index,
y=strat.equity_curve / strat.equity_curve.iloc[0] * 100,
name=strat.name,
line=dict(color=color),
), row=1, col=1)
fig.add_trace(go.Scatter(
x=[abs(strat.max_drawdown_pct)],
y=[strat.annual_return_pct],
mode='markers+text',
marker=dict(size=12, color=color),
text=[strat.name],
textposition='top center',
showlegend=False,
), row=1, col=2)
rolling_max = strat.equity_curve.cummax()
drawdown = (strat.equity_curve - rolling_max) / rolling_max * 100
fig.add_trace(go.Scatter(
x=drawdown.index,
y=drawdown,
fill='tozeroy',
name=strat.name,
line=dict(color=color),
showlegend=False,
opacity=0.6,
), row=2, col=2)
fig.update_layout(
title='Strategy Comparison Report',
height=900,
template='plotly_dark',
)
return fig
Visualization is especially useful when presenting results to a team or investors.
How Comparison is Performed: Step-by-Step
- Load historical data and a list of strategy configurations.
- Run backtesting — each strategy is executed on a unified dataset.
- The system collects equity curves and trade lists, then computes all metrics.
- The comparator creates a summary table, ranks strategies, and builds a correlation matrix.
- The report is exported in HTML with Plotly charts.
The entire process takes from a few minutes to an hour depending on the number of strategies and data volume. In practice, for 20 strategies on a one-year dataset, calculation takes under 5 minutes.
Example Summary for Three Strategies
| Strategy | Annual Return % | Sharpe Ratio | Max Drawdown % | Profit Factor |
|---|---|---|---|---|
| Momentum | 34.2 | 1.85 | -18.3 | 2.10 |
| MeanRev | 18.7 | 1.12 | -25.1 | 1.45 |
| Breakout | 41.5 | 2.10 | -22.7 | 2.55 |
From the table, Breakout shows the best Sharpe and return, but Momentum has lower drawdown. The ranking system, considering all metrics, will favor Breakout but also indicate high correlation between it and Momentum (if any).
How Portfolio Optimization Improves Sharpe?
If strategies are weakly correlated, they can be combined into a portfolio. We implemented optimization for Sharpe ratio:
def analyze_portfolio_combination(
strategy_returns: dict[str, pd.Series],
target_sharpe: float = 2.0,
) -> dict:
from scipy.optimize import minimize
returns_df = pd.DataFrame(strategy_returns).dropna()
mean_returns = returns_df.mean()
cov_matrix = returns_df.cov()
def neg_sharpe(weights):
portfolio_return = np.dot(weights, mean_returns) * 252
portfolio_vol = np.sqrt(np.dot(weights.T, np.dot(cov_matrix * 252, weights)))
return -portfolio_return / portfolio_vol if portfolio_vol > 0 else 999
n = len(strategy_returns)
constraints = [{'type': 'eq', 'fun': lambda w: np.sum(w) - 1}]
bounds = [(0, 1)] * n
x0 = [1/n] * n
result = minimize(neg_sharpe, x0, method='SLSQP', bounds=bounds, constraints=constraints)
optimal_weights = dict(zip(strategy_returns.keys(), result.x))
combined_returns = sum(ret * w for ret, w in zip(returns_df.values.T, result.x))
combined_sharpe = -result.fun
return {
'optimal_weights': optimal_weights,
'combined_sharpe': combined_sharpe,
'improvement_vs_best_single': combined_sharpe - max(
ret.mean() / ret.std() * np.sqrt(252) for ret in strategy_returns.values()
),
}
In practice, a combination of two to three uncorrelated strategies improves Sharpe by 30–50% relative to the best single strategy. Our experience developing such systems spans more than 50 projects.
What's Included in the Development of a Comparison System?
We provide:
- Architecture and design of the comparison module
- Implementation in Python using pandas, numpy, scipy
- Integration with your existing backtester
- Report generation as a DataFrame or HTML
- Visualization via Plotly
- Documentation and usage examples
- Post-deployment support
The development timeline for a turnkey system ranges from 2 to 4 weeks, and the cost is calculated individually based on integration complexity. The new system saves up to 80% of the time spent on manual analysis. We guarantee that all calculations comply with academic standards.
If you want an objective comparison system for your strategies, order development tailored to your requirements. Contact us for a consultation: describe your task, and we will propose a solution.
Typical Mistakes in Strategy Comparison
- Look-ahead bias: using future data when calculating metrics. We strictly separate in-sample and out-of-sample data.
- Survivor bias: ignoring dead strategies. We include all tested strategies.
- Ignoring commissions: commission drag can eat up to 2% annually. Our backtester accounts for them.
- Comparing over different periods: all metrics are annualized, so periods can be different.







