This notebook applies PCARVFLSimulator
(from the Python package synthe) — a GAN-like
tabular data synthesizer built from PCA scores, a Random Vector Functional-Link (RVFL) network,
and residual bootstrapping — to the French Motor Third-Party Liability (freMTPL2) data found in
PNM0792/auto-insurance-pricing,
and compares it against CTGAN, a GAN architecture purpose-built
for tabular data.
Data used (automobile/data/):
freMTPL2freq.csv— one row per policy (frequency file):IDpol, ClaimNb, Exposure, Area, VehPower, VehAge, DrivAge, BonusMalus, VehBrand, VehGas, Density, RegionfreMTPL2sev.csv— one row per claim (severity file):IDpol, ClaimAmount. Since this file alone is only 2 columns, we merge it back onto the policy’s rating factors from the frequency file (IDpoljoin) to get a proper multivariate severity dataset — mirroring what the repo’s ownsev_model.ipynbdoes internally.
What this notebook does, in 4 parts:
- Setup & data loading
freMTPL2freq— numeric features only: PCARVFLSimulator vs CTGANfreMTPL2sev(merged with policy features) — numeric features only: PCARVFLSimulator vs CTGAN- Both datasets again, this time encoding the categorical columns (
Area,VehBrand,VehGas,Region) into numeric features withcategory_encoders.TargetEncoder, then re-running the comparison
Each part fits both simulators on an identical sample, generates synthetic rows, and scores them
with synthe.adequacy_report() (MMD, energy distance, per-feature KS/AD tests, moment matching,
and a random-projection KS sweep), plus a marginal-distribution plot.
0. Setup
!pip install synthe ctgan category_encoders
# !pip install synthe ctgan category_encoders pandas numpy matplotlib scikit-learn optuna
import time
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from synthe import PCARVFLSimulator, adequacy_report
from ctgan import CTGAN
import category_encoders as ce
np.random.seed(42)
RANDOM_STATE = 42
N_TRAIN = 3000 # sub-sample size (full freq file has 678,013 rows)
N_SYN = 400 # synthetic rows generated, matches the blog post's convention
CTGAN_EPOCHS = 300
# Clone (or point to a local copy of) the data repo
# git clone https://github.com/PNM0792/auto-insurance-pricing.git
DATA_DIR = "auto-insurance-pricing/automobile/data"
freq = pd.read_csv("https://raw.githubusercontent.com/PNM0792/auto-insurance-pricing/refs/heads/main/automobile/data/freMTPL2freq.csv")
sev = pd.read_csv("https://raw.githubusercontent.com/PNM0792/auto-insurance-pricing/refs/heads/main/automobile/data/freMTPL2sev.csv")
print("freq:", freq.shape)
print("sev :", sev.shape)
freq.head()
freq: (678013, 12)
sev : (26639, 2)
| IDpol | ClaimNb | Exposure | Area | VehPower | VehAge | DrivAge | BonusMalus | VehBrand | VehGas | Density | Region | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1.00 | 1 | 0.10 | D | 5 | 0 | 55 | 50 | B12 | Regular | 1217 | R82 |
| 1 | 3.00 | 1 | 0.77 | D | 5 | 0 | 55 | 50 | B12 | Regular | 1217 | R82 |
| 2 | 5.00 | 1 | 0.75 | B | 6 | 2 | 52 | 50 | B12 | Diesel | 54 | R22 |
| 3 | 10.00 | 1 | 0.09 | B | 7 | 0 | 46 | 50 | B12 | Diesel | 76 | R72 |
| 4 | 11.00 | 1 | 0.84 | B | 7 | 0 | 46 | 50 | B12 | Diesel | 76 | R72 |
def plot_marginals(X_real, X_pcarvfl, X_ctgan, col_names, title, ncols=5):
n = len(col_names)
nrows = int(np.ceil(n / ncols))
fig, axes = plt.subplots(nrows, ncols, figsize=(4 * ncols, 4 * nrows))
axes = np.array(axes).ravel()
for i, c in enumerate(col_names):
ax = axes[i]
ax.hist(X_real[:, i], bins=30, density=True, alpha=0.5, label="Real", color="black")
ax.hist(X_pcarvfl[:, i], bins=30, density=True, alpha=0.5, label="PCARVFLSimulator", color="tab:blue")
ax.hist(X_ctgan[:, i], bins=30, density=True, alpha=0.5, label="CTGAN", color="tab:red")
ax.set_title(c, fontsize=10)
if i == 0:
ax.legend(fontsize=8)
for j in range(n, len(axes)):
axes[j].axis("off")
plt.suptitle(title, fontsize=13)
plt.tight_layout()
plt.show()
1. freMTPL2freq — numeric features only
Features: Exposure, VehPower, VehAge, DrivAge, BonusMalus, Density (log1p-transformed,
since raw density spans 1–27,000 and is heavily right-skewed).
num_cols_freq = ["Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus", "Density"]
df_s = freq.sample(n=N_TRAIN, random_state=RANDOM_STATE).reset_index(drop=True)
X_df_freq = df_s[num_cols_freq].copy()
X_df_freq["Density"] = np.log1p(X_df_freq["Density"])
X_real_freq = X_df_freq.values.astype(float)
print("Data shape:", X_real_freq.shape)
Data shape: (3000, 6)
t0 = time.time()
sim_freq = PCARVFLSimulator(random_state=RANDOM_STATE)
sim_freq.fit(X_real_freq, n_trials=30)
fit_time_pcarvfl = time.time() - t0
X_syn_pcarvfl_freq = sim_freq.sample(N_SYN)
print(f"PCARVFLSimulator fit time: {fit_time_pcarvfl:.1f}s")
print("\n" + "=" * 56)
print(f" FREMTPL2FREQ NUMERIC (n={X_real_freq.shape[0]}, d={X_real_freq.shape[1]})")
print("=" * 56)
report_pcarvfl_freq = adequacy_report(X_real_freq, X_syn_pcarvfl_freq)
[PCARVFL] 6 components (100.0% var)
[PCARVFL] nodes=153 α=5.06e+00 mmd=0.00000
PCARVFLSimulator fit time: 16.0s
========================================================
FREMTPL2FREQ NUMERIC (n=3000, d=6)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=6)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00088
Energy distance 0.02645
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.13514
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 0.667 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 5.61936
AD p Bonferroni ↑ 0.00600 ✗ reject
AD reject rate ↓ 0.667 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.32427
Std MAE ↓ 0.12301
Std ratio 0.9957 (want ≈ 1.00)
Corr Frobenius ↓ 0.2866
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.05163 over 50 directions
────────────────────────────────────────────────────────
t0 = time.time()
model_freq = CTGAN(epochs=CTGAN_EPOCHS, cuda=False)
model_freq.fit(X_df_freq)
fit_time_ctgan = time.time() - t0
print(f"CTGAN fit time: {fit_time_ctgan:.1f}s")
X_syn_ctgan_freq = model_freq.sample(N_SYN).values.astype(float)
print("\n" + "=" * 56)
print(f" FREMTPL2FREQ NUMERIC - CTGAN (n={X_real_freq.shape[0]}, d={X_real_freq.shape[1]})")
print("=" * 56)
report_ctgan_freq = adequacy_report(X_real_freq, X_syn_ctgan_freq)
CTGAN fit time: 116.5s
========================================================
FREMTPL2FREQ NUMERIC - CTGAN (n=3000, d=6)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=6)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00083
Energy distance 0.09829
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.16719
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 1.000 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 28.65123
AD p Bonferroni ↑ 0.00600 ✗ reject
AD reject rate ↓ 1.000 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 1.15431
Std MAE ↓ 0.56269
Std ratio 0.9197 (want ≈ 1.00)
Corr Frobenius ↓ 0.8853
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.10330 over 50 directions
────────────────────────────────────────────────────────
plot_marginals(
X_real_freq, X_syn_pcarvfl_freq, X_syn_ctgan_freq,
col_names=["Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus", "log1p(Density)"],
title=f"freMTPL2freq (n={N_TRAIN} sample): Real vs Synthetic Marginals",
)

2. freMTPL2sev (merged with policy features) — numeric features only
The severity file alone is just IDpol, ClaimAmount, so we merge it back onto the frequency
file’s rating factors to build a proper multivariate severity dataset. Features:
ClaimAmount (log1p), Exposure, VehPower, VehAge, DrivAge, BonusMalus, Density (log1p).
merged = sev.merge(freq, on="IDpol", how="left").dropna().reset_index(drop=True)
print("Merged severity dataset:", merged.shape)
num_cols_sev = ["ClaimAmount", "Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus", "Density"]
df_s_sev = merged.sample(n=N_TRAIN, random_state=RANDOM_STATE).reset_index(drop=True)
X_df_sev = df_s_sev[num_cols_sev].copy()
X_df_sev["ClaimAmount"] = np.log1p(X_df_sev["ClaimAmount"])
X_df_sev["Density"] = np.log1p(X_df_sev["Density"])
X_real_sev = X_df_sev.values.astype(float)
print("Data shape:", X_real_sev.shape)
Merged severity dataset: (26444, 13)
Data shape: (3000, 7)
t0 = time.time()
sim_sev = PCARVFLSimulator(random_state=RANDOM_STATE)
sim_sev.fit(X_real_sev, n_trials=30)
fit_time_pcarvfl_sev = time.time() - t0
X_syn_pcarvfl_sev = sim_sev.sample(N_SYN)
print(f"PCARVFLSimulator fit time: {fit_time_pcarvfl_sev:.1f}s")
print("\n" + "=" * 56)
print(f" FREMTPL2SEV (merged w/ policy features) (n={X_real_sev.shape[0]}, d={X_real_sev.shape[1]})")
print("=" * 56)
report_pcarvfl_sev = adequacy_report(X_real_sev, X_syn_pcarvfl_sev)
[PCARVFL] 7 components (100.0% var)
[PCARVFL] nodes=605 α=1.88e-04 mmd=0.00000
PCARVFLSimulator fit time: 15.5s
========================================================
FREMTPL2SEV (merged w/ policy features) (n=3000, d=7)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=7)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00076
Energy distance 0.01769
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.13814
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 0.857 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 3.81354
AD p Bonferroni ↑ 0.00700 ✗ reject
AD reject rate ↓ 0.571 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.41567
Std MAE ↓ 0.16528
Std ratio 0.9817 (want ≈ 1.00)
Corr Frobenius ↓ 0.3065
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.05793 over 50 directions
────────────────────────────────────────────────────────
t0 = time.time()
model_sev = CTGAN(epochs=CTGAN_EPOCHS, cuda=False)
model_sev.fit(X_df_sev)
fit_time_ctgan_sev = time.time() - t0
print(f"CTGAN fit time: {fit_time_ctgan_sev:.1f}s")
X_syn_ctgan_sev = model_sev.sample(N_SYN).values.astype(float)
print("\n" + "=" * 56)
print(f" FREMTPL2SEV - CTGAN (n={X_real_sev.shape[0]}, d={X_real_sev.shape[1]})")
print("=" * 56)
report_ctgan_sev = adequacy_report(X_real_sev, X_syn_ctgan_sev)
CTGAN fit time: 106.5s
========================================================
FREMTPL2SEV - CTGAN (n=3000, d=7)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=7)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00000
Energy distance 0.19864
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.26067
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 0.857 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 46.46704
AD p Bonferroni ↑ 0.00700 ✗ reject
AD reject rate ↓ 0.857 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 2.13144
Std MAE ↓ 0.76668
Std ratio 0.9633 (want ≈ 1.00)
Corr Frobenius ↓ 0.9285
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.14216 over 50 directions
────────────────────────────────────────────────────────
plot_marginals(
X_real_sev, X_syn_pcarvfl_sev, X_syn_ctgan_sev,
col_names=["log1p(ClaimAmount)", "Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus", "log1p(Density)"],
title=f"freMTPL2sev merged w/ policy features (n={N_TRAIN} sample): Real vs Synthetic Marginals",
ncols=4,
)

3. Adding the categorical features with category_encoders
Area, VehBrand, VehGas, Region are text columns. We convert them to numeric using
category_encoders.TargetEncoder, which is a standard actuarial technique — it replaces each
category with a smoothed mean of the response for that category (a “relativity”):
- for the frequency dataset, we target-encode against
ClaimNb - for the severity dataset, we target-encode against
ClaimAmount
3a. freMTPL2freq — numeric + target-encoded categoricals
cat_cols = ["Area", "VehBrand", "VehGas", "Region"]
encoder_freq = ce.TargetEncoder(cols=cat_cols, smoothing=10.0)
X_cat_enc_freq = encoder_freq.fit_transform(df_s[cat_cols], df_s["ClaimNb"])
X_cat_enc_freq.columns = [f"{c}_te" for c in cat_cols]
X_df_freq_full = pd.concat([df_s[num_cols_freq].copy(), X_cat_enc_freq], axis=1)
X_df_freq_full["Density"] = np.log1p(X_df_freq_full["Density"])
X_real_freq_full = X_df_freq_full.values.astype(float)
print("Data shape:", X_real_freq_full.shape)
print("Features:", list(X_df_freq_full.columns))
Data shape: (3000, 10)
Features: ['Exposure', 'VehPower', 'VehAge', 'DrivAge', 'BonusMalus', 'Density', 'Area_te', 'VehBrand_te', 'VehGas_te', 'Region_te']
t0 = time.time()
sim_freq_full = PCARVFLSimulator(random_state=RANDOM_STATE)
sim_freq_full.fit(X_real_freq_full, n_trials=30)
fit_time_pcarvfl_freq_full = time.time() - t0
X_syn_pcarvfl_freq_full = sim_freq_full.sample(N_SYN)
print(f"PCARVFLSimulator fit time: {fit_time_pcarvfl_freq_full:.1f}s")
print("\n" + "=" * 56)
print(f" FREMTPL2FREQ FULL (num+target-enc cat) (n={X_real_freq_full.shape[0]}, d={X_real_freq_full.shape[1]})")
print("=" * 56)
report_pcarvfl_freq_full = adequacy_report(X_real_freq_full, X_syn_pcarvfl_freq_full)
[PCARVFL] 9 components (95.1% var)
[PCARVFL] nodes=365 α=3.49e-03 mmd=0.00017
PCARVFLSimulator fit time: 17.1s
========================================================
FREMTPL2FREQ FULL (num+target-enc cat) (n=3000, d=10)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=10)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00210
Energy distance 0.03199
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.16247
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 0.800 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 8.79878
AD p Bonferroni ↑ 0.01000 ✗ reject
AD reject rate ↓ 0.800 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.19352
Std MAE ↓ 0.08292
Std ratio 0.9892 (want ≈ 1.00)
Corr Frobenius ↓ 0.5270
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.04996 over 50 directions
────────────────────────────────────────────────────────
t0 = time.time()
model_freq_full = CTGAN(epochs=CTGAN_EPOCHS, cuda=False)
model_freq_full.fit(X_df_freq_full)
fit_time_ctgan_freq_full = time.time() - t0
print(f"CTGAN fit time: {fit_time_ctgan_freq_full:.1f}s")
X_syn_ctgan_freq_full = model_freq_full.sample(N_SYN).values.astype(float)
print("\n" + "=" * 56)
print(f" FREMTPL2FREQ FULL - CTGAN (n={X_real_freq_full.shape[0]}, d={X_real_freq_full.shape[1]})")
print("=" * 56)
report_ctgan_freq_full = adequacy_report(X_real_freq_full, X_syn_ctgan_freq_full)
CTGAN fit time: 121.9s
========================================================
FREMTPL2FREQ FULL - CTGAN (n=3000, d=10)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=10)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00000
Energy distance 0.25518
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.20975
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 1.000 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 35.95134
AD p Bonferroni ↑ 0.01000 ✗ reject
AD reject rate ↓ 1.000 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.70767
Std MAE ↓ 0.30220
Std ratio 1.0647 (want ≈ 1.00)
Corr Frobenius ↓ 1.1302
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.13465 over 50 directions
────────────────────────────────────────────────────────
plot_marginals(
X_real_freq_full, X_syn_pcarvfl_freq_full, X_syn_ctgan_freq_full,
col_names=["Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus", "log1p(Density)",
"Area_te", "VehBrand_te", "VehGas_te", "Region_te"],
title=f"freMTPL2freq FULL (num + target-encoded cat, n={N_TRAIN}): Real vs Synthetic Marginals",
)

3b. freMTPL2sev (merged) — numeric + target-encoded categoricals
encoder_sev = ce.TargetEncoder(cols=cat_cols, smoothing=10.0)
X_cat_enc_sev = encoder_sev.fit_transform(df_s_sev[cat_cols], df_s_sev["ClaimAmount"])
X_cat_enc_sev.columns = [f"{c}_te" for c in cat_cols]
X_df_sev_full = pd.concat([df_s_sev[["ClaimAmount"]].copy(), df_s_sev[num_cols_sev[1:]].copy(), X_cat_enc_sev], axis=1)
X_df_sev_full["ClaimAmount"] = np.log1p(X_df_sev_full["ClaimAmount"])
X_df_sev_full["Density"] = np.log1p(X_df_sev_full["Density"])
for c in X_cat_enc_sev.columns:
X_df_sev_full[c] = np.log1p(X_df_sev_full[c].clip(lower=0))
X_real_sev_full = X_df_sev_full.values.astype(float)
print("Data shape:", X_real_sev_full.shape)
print("Features:", list(X_df_sev_full.columns))
Data shape: (3000, 11)
Features: ['ClaimAmount', 'Exposure', 'VehPower', 'VehAge', 'DrivAge', 'BonusMalus', 'Density', 'Area_te', 'VehBrand_te', 'VehGas_te', 'Region_te']
t0 = time.time()
sim_sev_full = PCARVFLSimulator(random_state=RANDOM_STATE)
sim_sev_full.fit(X_real_sev_full, n_trials=30)
fit_time_pcarvfl_sev_full = time.time() - t0
X_syn_pcarvfl_sev_full = sim_sev_full.sample(N_SYN)
print(f"PCARVFLSimulator fit time: {fit_time_pcarvfl_sev_full:.1f}s")
print("\n" + "=" * 56)
print(f" FREMTPL2SEV FULL (num+target-enc cat) (n={X_real_sev_full.shape[0]}, d={X_real_sev_full.shape[1]})")
print("=" * 56)
report_pcarvfl_sev_full = adequacy_report(X_real_sev_full, X_syn_pcarvfl_sev_full)
[PCARVFL] 10 components (97.5% var)
[PCARVFL] nodes=181 α=5.59e-04 mmd=0.00032
PCARVFLSimulator fit time: 13.4s
========================================================
FREMTPL2SEV FULL (num+target-enc cat) (n=3000, d=11)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=11)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00113
Energy distance 0.02130
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.15955
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 0.909 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 8.03864
AD p Bonferroni ↑ 0.01100 ✗ reject
AD reject rate ↓ 0.636 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.26444
Std MAE ↓ 0.11333
Std ratio 0.9822 (want ≈ 1.00)
Corr Frobenius ↓ 0.5038
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.05155 over 50 directions
────────────────────────────────────────────────────────
t0 = time.time()
model_sev_full = CTGAN(epochs=CTGAN_EPOCHS, cuda=False)
model_sev_full.fit(X_df_sev_full)
fit_time_ctgan_sev_full = time.time() - t0
print(f"CTGAN fit time: {fit_time_ctgan_sev_full:.1f}s")
X_syn_ctgan_sev_full = model_sev_full.sample(N_SYN).values.astype(float)
print("\n" + "=" * 56)
print(f" FREMTPL2SEV FULL - CTGAN (n={X_real_sev_full.shape[0]}, d={X_real_sev_full.shape[1]})")
print("=" * 56)
report_ctgan_sev_full = adequacy_report(X_real_sev_full, X_syn_ctgan_sev_full)
CTGAN fit time: 127.2s
========================================================
FREMTPL2SEV FULL - CTGAN (n=3000, d=11)
========================================================
────────────────────────────────────────────────────────
Adequacy Report (n_real=3000, n_syn=400, d=11)
────────────────────────────────────────────────────────
── Distributional distance [standardised space]
MMD (biased, ≥0) 0.00000
Energy distance 0.19999
── Per-feature KS tests [original scale]
KS stat (mean ↓) 0.25571
KS p Bonferroni ↑ 0.00000 ✗ reject
KS reject rate ↓ 1.000 (α=0.05)
── Per-feature AD tests [original scale]
AD stat (mean ↓) 41.12033
AD p Bonferroni ↑ 0.01100 ✗ reject
AD reject rate ↓ 1.000 (α=0.05)
── Moment matching [original scale]
Mean MAE ↓ 0.93624
Std MAE ↓ 0.21834
Std ratio 0.9583 (want ≈ 1.00)
Corr Frobenius ↓ 1.5596
── Random-projection sweep [standardised space]
KS proj (mean ↓) 0.10638 over 50 directions
────────────────────────────────────────────────────────
plot_marginals(
X_real_sev_full, X_syn_pcarvfl_sev_full, X_syn_ctgan_sev_full,
col_names=["log1p(ClaimAmount)", "Exposure", "VehPower", "VehAge", "DrivAge", "BonusMalus",
"log1p(Density)", "log1p(Area_te)", "log1p(VehBrand_te)", "log1p(VehGas_te)", "log1p(Region_te)"],
title=f"freMTPL2sev FULL (num + target-encoded cat, n={N_TRAIN}): Real vs Synthetic Marginals",
ncols=4,
)

4. Summary tables
Recap of the adequacy_report() output across all four runs (numbers below are from the
run performed while writing this notebook; re-running with RANDOM_STATE=42 should reproduce
them closely, modulo Optuna/CTGAN’s own internal stochasticity).
freMTPL2freq — numeric only
| Metric | Direction | PCARVFLSimulator | CTGAN (300 epochs) |
|---|---|---|---|
| Fit time | — | 4.9s | 56.1s |
| MMD (biased) | ↓ better | 0.00088 | 0.00169 |
| Energy distance | ↓ better | 0.02645 | 0.17676 |
| KS stat (mean) | ↓ better | 0.13514 | 0.15433 |
| KS reject rate | ↓ better | 0.667 | 0.833 |
| AD stat (mean) | ↓ better | 5.61936 | 34.60157 |
| AD reject rate | ↓ better | 0.667 | 0.833 |
| Mean MAE | ↓ better | 0.32427 | 1.42774 |
| Std MAE | ↓ better | 0.12301 | 1.17582 |
| Std ratio (want ≈1) | closer to 1 | 0.9957 | 0.8222 |
| Corr Frobenius | ↓ better | 0.2866 | 0.7769 |
| KS proj (50 dirs) | ↓ better | 0.05163 | 0.12971 |
freMTPL2sev (merged) — numeric only
| Metric | Direction | PCARVFLSimulator | CTGAN (300 epochs) |
|---|---|---|---|
| Fit time | — | 3.8s | 59.0s |
| MMD (biased) | ↓ better | 0.00076 | 0.00018 |
| Energy distance | ↓ better | 0.01769 | 0.06169 |
| KS stat (mean) | ↓ better | 0.13814 | 0.19414 |
| KS reject rate | ↓ better | 0.857 | 0.857 |
| AD stat (mean) | ↓ better | 3.81354 | 25.92328 |
| AD reject rate | ↓ better | 0.571 | 0.857 |
| Mean MAE | ↓ better | 0.41567 | 0.86539 |
| Std MAE | ↓ better | 0.16528 | 0.47315 |
| Std ratio (want ≈1) | closer to 1 | 0.9817 | 0.9701 |
| Corr Frobenius | ↓ better | 0.3065 | 0.8698 |
| KS proj (50 dirs) | ↓ better | 0.05793 | 0.08365 |
freMTPL2freq — numeric + target-encoded categoricals
| Metric | Direction | PCARVFLSimulator | CTGAN (300 epochs) |
|---|---|---|---|
| Fit time | — | 4.3s | 54.8s |
| MMD (biased) | ↓ better | 0.00210 | 0.00000 |
| Energy distance | ↓ better | 0.03199 | 0.32201 |
| KS stat (mean) | ↓ better | 0.16247 | 0.21083 |
| KS reject rate | ↓ better | 0.800 | 0.900 |
| AD stat (mean) | ↓ better | 8.79878 | 63.19130 |
| AD reject rate | ↓ better | 0.800 | 1.000 |
| Mean MAE | ↓ better | 0.19352 | 0.96904 |
| Std MAE | ↓ better | 0.08292 | 0.66499 |
| Std ratio (want ≈1) | closer to 1 | 0.9892 | 0.8339 |
| Corr Frobenius | ↓ better | 0.5270 | 1.0160 |
| KS proj (50 dirs) | ↓ better | 0.04996 | 0.13840 |
freMTPL2sev (merged) — numeric + target-encoded categoricals
| Metric | Direction | PCARVFLSimulator | CTGAN (300 epochs) |
|---|---|---|---|
| Fit time | — | 4.0s | 57.7s |
| MMD (biased) | ↓ better | 0.00113 | 0.00071 |
| Energy distance | ↓ better | 0.02130 | 0.13080 |
| KS stat (mean) | ↓ better | 0.15955 | 0.20471 |
| KS reject rate | ↓ better | 0.909 | 1.000 |
| AD stat (mean) | ↓ better | 8.03864 | 20.63413 |
| AD reject rate | ↓ better | 0.636 | 1.000 |
| Mean MAE | ↓ better | 0.26444 | 0.47455 |
| Std MAE | ↓ better | 0.11333 | 0.69187 |
| Std ratio (want ≈1) | closer to 1 | 0.9822 | 1.0198 |
| Corr Frobenius | ↓ better | 0.5038 | 1.5450 |
| KS proj (50 dirs) | ↓ better | 0.05155 | 0.10660 |
Across all four runs, PCARVFLSimulator wins on the large majority of adequacy metrics — energy distance, moment matching, correlation-structure preservation, and the random-projection KS sweep — while fitting roughly 10–15x faster than CTGAN (seconds vs ~1 minute at this sample size; the gap would grow further on the full 678k-row file).
For attribution, please cite this work as:
T. Moudiki (2026-08-22). PCARVFL vs CTGAN for synthetic tabular data generation on an insurance pricing dataset. Retrieved from https://thierrymoudiki.github.io/blog/2026/08/22/python/pcarvfl-vs-ctgan-insurance
BibTeX citation (remove empty spaces)
@misc{ tmoudiki20260822,
author = { T. Moudiki },
title = { PCARVFL vs CTGAN for synthetic tabular data generation on an insurance pricing dataset },
url = { https://thierrymoudiki.github.io/blog/2026/08/22/python/pcarvfl-vs-ctgan-insurance },
year = { 2026 } }
Previous publications
- PCARVFL vs CTGAN for synthetic tabular data generation on an insurance pricing dataset Aug 22, 2026
- 'Zero-Shot Probabilistic Stock Returns Forecasting with Pretrained RVFL Networks' accepted at COPA 2026 (and to appear in the Proceedings of Machine Learning Research) Aug 15, 2026
- 'PCARVFLSimulator': a GAN-like tabular data synthesizer built from PCA scores, a Random Vector Functional-Link network, and residuals bootstrapping Aug 10, 2026
- 'garchf': GARCH probabilistic forecasting with package 'forecast'-style interface (and 'rugarch' under the hood) Aug 1, 2026
- GPopt for R: Bayesian and conformal optimization of black-box functions and hyperparameter tuning Jul 26, 2026
- My last R posts: How conformalization helps weak models, fast conformal prediction with jackknife+ (and no refitting), and sklearn in R Jul 13, 2026
- Natively Interpretable Boosting Jul 12, 2026
- Fast conformal prediction (no refitting) for some Machine Learning models via closed-form jackknife plus Jun 27, 2026
- Using scikit-learn models in R easily with the tisthemachinelearner package Jun 21, 2026
- No-Code Machine Learning in Excel with the Techtonique API Jun 14, 2026
- How Conformal Prediction Makes Linear Models Good Enough — An Example Using R Package mlS3 Jun 7, 2026
- Techtonique dot net, the Machine Learning web API, is back online (but more like a passion project for now) May 31, 2026
- Conformalized TabICL: Prediction Intervals for a State-Of-The-Art Tabular Foundation Model in Python and R May 21, 2026
- Conformalized TabPFN: Prediction Intervals for a Pretrained Transformer for Tabular Data in Python and R May 17, 2026
- Probabilistic Time Series Cross-Validation with R package crossvalidation May 16, 2026
- One interface, (Almost) Every Classifier (and Regressor): unifiedml v0.3.0 May 9, 2026
- You Don't Need to Learn All the Weights on tabular data: The Case for rvflnet (a nonlinear expressive glmnet) on regression, classification and survival analysis May 2, 2026
- Survival analysis with sklearn, glmnet, keras, pytorch, lightgbm, xgboost, nnetsauce, mlsauce Part 2 Apr 28, 2026
- Any Sklearn Regressor as a Survival Model — Does It Actually Work? Benchmarking vs Established Packages Apr 26, 2026
- Conformal Optimization Beats Bayesian Optimization, Optuna and Random Search on 72 classification Datasets Apr 19, 2026
- `mlS3` — A Unified S3 Machine Learning Interface in R Apr 12, 2026
- One interface, (Almost) Every Classifier: unifiedml v0.2.1 Apr 4, 2026
- Techtonique dot net is down until further notice Apr 1, 2026
- Explaining Time-Series Forecasts with Sensitivity Analysis (ahead::dynrmf and external regressors) Mar 29, 2026
- Python version of 'Option pricing using time series models as market price of risk Pt.3' Mar 22, 2026
- Option pricing using time series models as market price of risk Pt.3 Mar 16, 2026
- Explaining Time-Series Forecasts with Exact Shapley Values (ahead::dynrmf with external regressors applied to scenarios) Mar 8, 2026
- My Presentation at Risk 2026: Lightweight Transfer Learning for Financial Forecasting Mar 1, 2026
- nnetsauce with and without jax for GPU acceleration Feb 23, 2026
- Understanding Boosted Configuration Networks (combined neural networks and boosting): An Intuitive Guide Through Their Hyperparameters Feb 16, 2026
- R version of Python package survivalist, for model-agnostic survival analysis Feb 9, 2026
- Presenting Lightweight Transfer Learning for Financial Forecasting (Risk 2026) Feb 4, 2026
- Option pricing using time series models as market price of risk Feb 1, 2026
- Enhancing Time Series Forecasting (ahead::ridge2f) with Attention-Based Context Vectors (ahead::contextridge2f) Jan 31, 2026
- Overfitting and scaling (on GPU T4) tests on nnetsauce.CustomRegressor Jan 29, 2026
- Beyond Cross-validation: Hyperparameter Optimization via Generalization Gap Modeling Jan 25, 2026
- GPopt for Machine Learning (hyperparameters' tuning) Jan 21, 2026
- rtopy: an R to Python bridge -- novelties Jan 8, 2026
- Python examples for 'Beyond Nelson-Siegel and splines: A model- agnostic Machine Learning framework for discount curve calibration, interpolation and extrapolation' Jan 3, 2026
- Forecasting benchmark: Dynrmf (a new serious competitor in town) vs Theta Method on M-Competitions and Tourism competitition Jan 1, 2026
- Finally figured out a way to port python packages to R using uv and reticulate: example with nnetsauce Dec 17, 2025
- Overfitting Random Fourier Features: Universal Approximation Property Dec 13, 2025
- Counterfactual Scenario Analysis with ahead::ridge2f Dec 11, 2025
- Zero-Shot Probabilistic Time Series Forecasting with TabPFN 2.5 and nnetsauce Dec 10, 2025
- ARIMA Pricing: Semi-Parametric Market price of risk for Risk-Neutral Pricing (code + preprint) Dec 7, 2025
- Analyzing Paper Reviews with LLMs: I Used ChatGPT, DeepSeek, Qwen, Mistral, Gemini, and Claude (and you should too + publish the analysis) Dec 3, 2025
- tisthemachinelearner: New Workflow with uv for R Integration of scikit-learn Dec 1, 2025
- (ICYMI) RPweave: Unified R + Python + LaTeX System using uv Nov 21, 2025
- unifiedml: A Unified Machine Learning Interface for R, is now on CRAN + Discussion about AI replacing humans Nov 16, 2025
- Context-aware Theta forecasting Method: Extending Classical Time Series Forecasting with Machine Learning Nov 13, 2025
- unifiedml in R: A Unified Machine Learning Interface Nov 5, 2025
- Deterministic Shift Adjustment in Arbitrage-Free Pricing (historical to risk-neutral short rates) Oct 28, 2025
- New instantaneous short rates models with their deterministic shift adjustment, for historical and risk-neutral simulation Oct 27, 2025
- RPweave: Unified R + Python + LaTeX System using uv Oct 19, 2025
- GAN-like Synthetic Data Generation Examples (on univariate, multivariate distributions, digits recognition, Fashion-MNIST, stock returns, and Olivetti faces) with DistroSimulator Oct 19, 2025
- R port of llama2.c Oct 9, 2025
- Native uncertainty quantification for time series with NGBoost Oct 8, 2025
- NGBoost (Natural Gradient Boosting) for Regression, Classification, Time Series forecasting and Reserving Oct 6, 2025
- Real-time pricing with a pretrained probabilistic stock return model Oct 1, 2025
- Combining any model with GARCH(1,1) for probabilistic stock forecasting Sep 23, 2025
- Generating Synthetic Data with R-vine Copulas using esgtoolkit in R Sep 21, 2025
- Reimagining Equity Solvency Capital Requirement Approximation (one of my Master's Thesis subjects): From Bilinear Interpolation to Probabilistic Machine Learning Sep 16, 2025
- Transfer Learning using ahead::ridge2f on synthetic stocks returns Pt.2: synthetic data generation Sep 9, 2025
- Transfer Learning using ahead::ridge2f on synthetic stocks returns Sep 8, 2025
- I'm supposed to present 'Conformal Predictive Simulations for Univariate Time Series' at COPA CONFERENCE 2025 in London... Sep 4, 2025
- external regressors in ahead::dynrmf's interface for Machine learning forecasting Sep 1, 2025
- Another interesting decision, now for 'Beyond Nelson-Siegel and splines: A model-agnostic Machine Learning framework for discount curve calibration, interpolation and extrapolation' Aug 20, 2025
- Boosting any randomized based learner for regression, classification and univariate/multivariate time series forcasting Jul 26, 2025
- New nnetsauce version with CustomBackPropRegressor (CustomRegressor with Backpropagation) and ElasticNet2Regressor (Ridge2 with ElasticNet regularization) Jul 15, 2025
- mlsauce (home to a model-agnostic gradient boosting algorithm) can now be installed from PyPI. Jul 10, 2025
- A user-friendly graphical interface to techtonique dot net's API (will eventually contain graphics). Jul 8, 2025
- Calling =TECHTO_MLCLASSIFICATION for Machine Learning supervised CLASSIFICATION in Excel is just a matter of copying and pasting Jul 7, 2025
- Calling =TECHTO_MLREGRESSION for Machine Learning supervised regression in Excel is just a matter of copying and pasting Jul 6, 2025
- Calling =TECHTO_RESERVING and =TECHTO_MLRESERVING for claims triangle reserving in Excel is just a matter of copying and pasting Jul 5, 2025
- Calling =TECHTO_SURVIVAL for Survival Analysis in Excel is just a matter of copying and pasting Jul 4, 2025
- Calling =TECHTO_SIMULATION for Stochastic Simulation in Excel is just a matter of copying and pasting Jul 3, 2025
- Calling =TECHTO_FORECAST for forecasting in Excel is just a matter of copying and pasting Jul 2, 2025
- Random Vector Functional Link (RVFL) artificial neural network with 2 regularization parameters successfully used for forecasting/synthetic simulation in professional settings: Extensions (including Bayesian) Jul 1, 2025
- R version of 'Backpropagating quasi-randomized neural networks' Jun 24, 2025
- Backpropagating quasi-randomized neural networks Jun 23, 2025
- Beyond ARMA-GARCH: leveraging any statistical model for volatility forecasting Jun 21, 2025
- Stacked generalization (Machine Learning model stacking) + conformal prediction for forecasting with ahead::mlf Jun 18, 2025
- An Overfitting dilemma: XGBoost Default Hyperparameters vs GenericBooster + LinearRegression Default Hyperparameters Jun 14, 2025
- Programming language-agnostic reserving using RidgeCV, LightGBM, XGBoost, and ExtraTrees Machine Learning models Jun 13, 2025
- Free R, Python and SQL editors in techtonique dot net Jun 9, 2025
- Beyond Nelson-Siegel and splines: A model-agnostic Machine Learning framework for discount curve calibration, interpolation and extrapolation Jun 7, 2025
- scikit-learn, glmnet, xgboost, lightgbm, pytorch, keras, nnetsauce in probabilistic Machine Learning (for longitudinal data) Reserving (work in progress) Jun 6, 2025
- R version of Probabilistic Machine Learning (for longitudinal data) Reserving (work in progress) Jun 5, 2025
- Probabilistic Machine Learning (for longitudinal data) Reserving (work in progress) Jun 4, 2025
- Python version of Beyond ARMA-GARCH: leveraging model-agnostic Quasi-Randomized networks and conformal prediction for nonparametric probabilistic stock forecasting (ML-ARCH) Jun 3, 2025
- Beyond ARMA-GARCH: leveraging model-agnostic Machine Learning and conformal prediction for nonparametric probabilistic stock forecasting (ML-ARCH) Jun 2, 2025
- Permutations and SHAPley values for feature importance in techtonique dot net's API (with R + Python + the command line) Jun 1, 2025
- Which patient is going to survive longer? Another guide to using techtonique dot net's API (with R + Python + the command line) for survival analysis May 31, 2025
- A Guide to Using techtonique.net's API and rush for simulating and plotting Stochastic Scenarios May 30, 2025
- Simulating Stochastic Scenarios with Diffusion Models: A Guide to Using techtonique.net's API for the purpose May 29, 2025
- Will my apartment in 5th avenue be overpriced or not? Harnessing the power of www.techtonique.net (+ xgboost, lightgbm, catboost) to find out May 28, 2025
- How long must I wait until something happens: A Comprehensive Guide to Survival Analysis via an API May 27, 2025
- Harnessing the Power of techtonique.net: A Comprehensive Guide to Machine Learning Classification via an API May 26, 2025
- Quantile regression with any regressor -- Examples with RandomForestRegressor, RidgeCV, KNeighborsRegressor May 20, 2025
- Survival stacking: survival analysis translated as supervised classification in R and Python May 5, 2025
- 'Bayesian' optimization of hyperparameters in a R machine learning model using the bayesianrvfl package Apr 25, 2025
- A lightweight interface to scikit-learn in R: Bayesian and Conformal prediction Apr 21, 2025
- A lightweight interface to scikit-learn in R Pt.2: probabilistic time series forecasting in conjunction with ahead::dynrmf Apr 20, 2025
- Extending the Theta forecasting method to GLMs, GAMs, GLMBOOST and attention: benchmarking on Tourism, M1, M3 and M4 competition data sets (28000 series) Apr 14, 2025
- Extending the Theta forecasting method to GLMs and attention Apr 8, 2025
- Nonlinear conformalized Generalized Linear Models (GLMs) with R package 'rvfl' (and other models) Mar 31, 2025
- Probabilistic Time Series Forecasting (predictive simulations) in Microsoft Excel using Python, xlwings lite and www.techtonique.net Mar 28, 2025
- Conformalize (improved prediction intervals and simulations) any R Machine Learning model with misc::conformalize Mar 25, 2025
- My poster for the 18th FINANCIAL RISKS INTERNATIONAL FORUM by Institut Louis Bachelier/Fondation du Risque/Europlace Institute of Finance Mar 19, 2025
- Interpretable probabilistic kernel ridge regression using Matérn 3/2 kernels Mar 16, 2025
- (News from) Probabilistic Forecasting of univariate and multivariate Time Series using Quasi-Randomized Neural Networks (Ridge2) and Conformal Prediction Mar 9, 2025
- Word-Online: re-creating Karpathy's char-RNN (with supervised linear online learning of word embeddings) for text completion Mar 8, 2025
- CRAN-like repository for most recent releases of Techtonique's R packages Mar 2, 2025
- Presenting 'Online Probabilistic Estimation of Carbon Beta and Carbon Shapley Values for Financial and Climate Risk' at Institut Louis Bachelier Feb 27, 2025
- Web app with DeepSeek R1 and Hugging Face API for chatting Feb 23, 2025
- tisthemachinelearner: A Lightweight interface to scikit-learn with 2 classes, Classifier and Regressor (in Python and R) Feb 17, 2025
- R version of survivalist: Probabilistic model-agnostic survival analysis using scikit-learn, xgboost, lightgbm (and conformal prediction) Feb 12, 2025
- Model-agnostic global Survival Prediction of Patients with Myeloid Leukemia in QRT/Gustave Roussy Challenge (challengedata.ens.fr): Python's survivalist Quickstart Feb 10, 2025
- A simple test of the martingale hypothesis in esgtoolkit Feb 3, 2025
- Command Line Interface (CLI) for techtonique.net's API Jan 31, 2025
- Gradient-Boosting and Boostrap aggregating anything (alert: high performance): Part5, easier install and Rust backend Jan 27, 2025
- Just got a paper on conformal prediction REJECTED by International Journal of Forecasting despite evidence on 30,000 time series (and more). What's going on? Part2: 1311 time series from the Tourism competition Jan 20, 2025
- Techtonique is released! (with a tutorial in various programming languages and formats) Jan 14, 2025
- Univariate and Multivariate Probabilistic Forecasting with nnetsauce and TabPFN Jan 14, 2025
- Just got a paper on conformal prediction REJECTED by International Journal of Forecasting despite evidence on 30,000 time series (and more). What's going on? Jan 5, 2025
- Python and Interactive dashboard version of Stock price forecasting with Deep Learning: throwing power at the problem (and why it won't make you rich) Dec 31, 2024
- Stock price forecasting with Deep Learning: throwing power at the problem (and why it won't make you rich) Dec 29, 2024
- No-code Machine Learning Cross-validation and Interpretability in techtonique.net Dec 23, 2024
- survivalist: Probabilistic model-agnostic survival analysis using scikit-learn, glmnet, xgboost, lightgbm, pytorch, keras, nnetsauce and mlsauce Dec 15, 2024
- Model-agnostic 'Bayesian' optimization (for hyperparameter tuning) using conformalized surrogates in GPopt Dec 9, 2024
- You can beat Forecasting LLMs (Large Language Models a.k.a foundation models) with nnetsauce.MTS Pt.2: Generic Gradient Boosting Dec 1, 2024
- You can beat Forecasting LLMs (Large Language Models a.k.a foundation models) with nnetsauce.MTS Nov 24, 2024
- Unified interface and conformal prediction (calibrated prediction intervals) for R package forecast (and 'affiliates') Nov 23, 2024
- GLMNet in Python: Generalized Linear Models Nov 18, 2024
- Gradient-Boosting anything (alert: high performance): Part4, Time series forecasting Nov 10, 2024
- Predictive scenarios simulation in R, Python and Excel using Techtonique API Nov 3, 2024
- Chat with your tabular data in www.techtonique.net Oct 30, 2024
- Gradient-Boosting anything (alert: high performance): Part3, Histogram-based boosting Oct 28, 2024
- R editor and SQL console (in addition to Python editors) in www.techtonique.net Oct 21, 2024
- R and Python consoles + JupyterLite in www.techtonique.net Oct 15, 2024
- Gradient-Boosting anything (alert: high performance): Part2, R version Oct 14, 2024
- Gradient-Boosting anything (alert: high performance) Oct 6, 2024
- Benchmarking 30 statistical/Machine Learning models on the VN1 Forecasting -- Accuracy challenge Oct 4, 2024
- Automated random variable distribution inference using Kullback-Leibler divergence and simulating best-fitting distribution Oct 2, 2024
- Forecasting in Excel using Techtonique's Machine Learning APIs under the hood Sep 30, 2024
- Techtonique web app for data-driven decisions using Mathematics, Statistics, Machine Learning, and Data Visualization Sep 25, 2024
- Parallel for loops (Map or Reduce) + New versions of nnetsauce and ahead Sep 16, 2024
- Adaptive (online/streaming) learning with uncertainty quantification using Polyak averaging in learningmachine Sep 10, 2024
- New versions of nnetsauce and ahead Sep 9, 2024
- Prediction sets and prediction intervals for conformalized Auto XGBoost, Auto LightGBM, Auto CatBoost, Auto GradientBoosting Sep 2, 2024
- Quick/automated R package development workflow (assuming you're using macOS or Linux) Part2 Aug 30, 2024
- R package development workflow (assuming you're using macOS or Linux) Aug 27, 2024
- A new method for deriving a nonparametric confidence interval for the mean Aug 26, 2024
- Conformalized adaptive (online/streaming) learning using learningmachine in Python and R Aug 19, 2024
- Bayesian (nonlinear) adaptive learning Aug 12, 2024
- Auto XGBoost, Auto LightGBM, Auto CatBoost, Auto GradientBoosting Aug 5, 2024
- Copulas for uncertainty quantification in time series forecasting Jul 28, 2024
- Forecasting uncertainty: sequential split conformal prediction + Block bootstrap (web app) Jul 22, 2024
- learningmachine for Python (new version) Jul 15, 2024
- learningmachine v2.0.0: Machine Learning with explanations and uncertainty quantification Jul 8, 2024
- My presentation at ISF 2024 conference (slides with nnetsauce probabilistic forecasting news) Jul 3, 2024
- 10 uncertainty quantification methods in nnetsauce forecasting Jul 1, 2024
- Forecasting with XGBoost embedded in Quasi-Randomized Neural Networks Jun 24, 2024
- Forecasting Monthly Airline Passenger Numbers with Quasi-Randomized Neural Networks Jun 17, 2024
- Automated hyperparameter tuning using any conformalized surrogate Jun 9, 2024
- Recognizing handwritten digits with Ridge2Classifier Jun 3, 2024
- Forecasting the Economy May 27, 2024
- A detailed introduction to Deep Quasi-Randomized 'neural' networks May 19, 2024
- Probability of receiving a loan; using learningmachine May 12, 2024
- mlsauce's `v0.18.2`: various examples and benchmarks with dimension reduction May 6, 2024
- mlsauce's `v0.17.0`: boosting with Elastic Net, polynomials and heterogeneity in explanatory variables Apr 29, 2024
- mlsauce's `v0.13.0`: taking into account inputs heterogeneity through clustering Apr 21, 2024
- mlsauce's `v0.12.0`: prediction intervals for LSBoostRegressor Apr 15, 2024
- Conformalized predictive simulations for univariate time series on more than 250 data sets Apr 7, 2024
- learningmachine v1.1.2: for Python Apr 1, 2024
- learningmachine v1.0.0: prediction intervals around the probability of the event 'a tumor being malignant' Mar 25, 2024
- Bayesian inference and conformal prediction (prediction intervals) in nnetsauce v0.18.1 Mar 18, 2024
- Multiple examples of Machine Learning forecasting with ahead Mar 11, 2024
- rtopy (v0.1.1): calling R functions in Python Mar 4, 2024
- ahead forecasting (v0.10.0): fast time series model calibration and Python plots Feb 26, 2024
- A plethora of datasets at your fingertips Part3: how many times do couples cheat on each other? Feb 19, 2024
- nnetsauce's introduction as of 2024-02-11 (new version 0.17.0) Feb 11, 2024
- Tuning Machine Learning models with GPopt's new version Part 2 Feb 5, 2024
- Tuning Machine Learning models with GPopt's new version Jan 29, 2024
- Subsampling continuous and discrete response variables Jan 22, 2024
- DeepMTS, a Deep Learning Model for Multivariate Time Series Jan 15, 2024
- A classifier that's very accurate (and deep) Pt.2: there are > 90 classifiers in nnetsauce Jan 8, 2024
- learningmachine: prediction intervals for conformalized Kernel ridge regression and Random Forest Jan 1, 2024
- A plethora of datasets at your fingertips Part2: how many times do couples cheat on each other? Descriptive analytics, interpretability and prediction intervals using conformal prediction Dec 25, 2023
- Diffusion models in Python with esgtoolkit (Part2) Dec 18, 2023
- Diffusion models in Python with esgtoolkit Dec 11, 2023
- Julia packaging at the command line Dec 4, 2023
- Quasi-randomized nnetworks in Julia, Python and R Nov 27, 2023
- A plethora of datasets at your fingertips Nov 20, 2023
- A classifier that's very accurate (and deep) Nov 12, 2023
- mlsauce version 0.8.10: Statistical/Machine Learning with Python and R Nov 5, 2023
- AutoML in nnetsauce (randomized and quasi-randomized nnetworks) Pt.2: multivariate time series forecasting Oct 29, 2023
- AutoML in nnetsauce (randomized and quasi-randomized nnetworks) Oct 22, 2023
- Version v0.14.0 of nnetsauce for R and Python Oct 16, 2023
- A diffusion model: G2++ Oct 9, 2023
- Diffusion models in ESGtoolkit + announcements Oct 2, 2023
- An infinity of time series forecasting models in nnetsauce (Part 2 with uncertainty quantification) Sep 25, 2023
- (News from) forecasting in Python with ahead (progress bars and plots) Sep 18, 2023
- Forecasting in Python with ahead Sep 11, 2023
- Risk-neutralize simulations Sep 4, 2023
- Comparing cross-validation results using crossval_ml and boxplots Aug 27, 2023
- Reminder Apr 30, 2023
- Did you ask ChatGPT about who you are? Apr 16, 2023
- A new version of nnetsauce (randomized and quasi-randomized 'neural' networks) Apr 2, 2023
- Simple interfaces to the forecasting API Nov 23, 2022
- A web application for forecasting in Python, R, Ruby, C#, JavaScript, PHP, Go, Rust, Java, MATLAB, etc. Nov 2, 2022
- Prediction intervals (not only) for Boosted Configuration Networks in Python Oct 5, 2022
- Boosted Configuration (neural) Networks Pt. 2 Sep 3, 2022
- Boosted Configuration (_neural_) Networks for classification Jul 21, 2022
- A Machine Learning workflow using Techtonique Jun 6, 2022
- Super Mario Bros © in the browser using PyScript May 8, 2022
- News from ESGtoolkit, ycinterextra, and nnetsauce Apr 4, 2022
- Explaining a Keras _neural_ network predictions with the-teller Mar 11, 2022
- New version of nnetsauce -- various quasi-randomized networks Feb 12, 2022
- A dashboard illustrating bivariate time series forecasting with `ahead` Jan 14, 2022
- Hundreds of Statistical/Machine Learning models for univariate time series, using ahead, ranger, xgboost, and caret Dec 20, 2021
- Forecasting with `ahead` (Python version) Dec 13, 2021
- Tuning and interpreting LSBoost Nov 15, 2021
- Time series cross-validation using `crossvalidation` (Part 2) Nov 7, 2021
- Fast and scalable forecasting with ahead::ridge2f Oct 31, 2021
- Automatic Forecasting with `ahead::dynrmf` and Ridge regression Oct 22, 2021
- Forecasting with `ahead` Oct 15, 2021
- Classification using linear regression Sep 26, 2021
- `crossvalidation` and random search for calibrating support vector machines Aug 6, 2021
- parallel grid search cross-validation using `crossvalidation` Jul 31, 2021
- `crossvalidation` on R-universe, plus a classification example Jul 23, 2021
- Documentation and source code for GPopt, a package for Bayesian optimization Jul 2, 2021
- Hyperparameters tuning with GPopt Jun 11, 2021
- A forecasting tool (API) with examples in curl, R, Python May 28, 2021
- Bayesian Optimization with GPopt Part 2 (save and resume) Apr 30, 2021
- Bayesian Optimization with GPopt Apr 16, 2021
- Compatibility of nnetsauce and mlsauce with scikit-learn Mar 26, 2021
- Explaining xgboost predictions with the teller Mar 12, 2021
- An infinity of time series models in nnetsauce Mar 6, 2021
- New activation functions in mlsauce's LSBoost Feb 12, 2021
- 2020 recap, Gradient Boosting, Generalized Linear Models, AdaOpt with nnetsauce and mlsauce Dec 29, 2020
- A deeper learning architecture in nnetsauce Dec 18, 2020
- Classify penguins with nnetsauce's MultitaskClassifier Dec 11, 2020
- Bayesian forecasting for uni/multivariate time series Dec 4, 2020
- Generalized nonlinear models in nnetsauce Nov 28, 2020
- Boosting nonlinear penalized least squares Nov 21, 2020
- Statistical/Machine Learning explainability using Kernel Ridge Regression surrogates Nov 6, 2020
- NEWS Oct 30, 2020
- A glimpse into my PhD journey Oct 23, 2020
- Submitting R package to CRAN Oct 16, 2020
- Simulation of dependent variables in ESGtoolkit Oct 9, 2020
- Forecasting lung disease progression Oct 2, 2020
- New nnetsauce Sep 25, 2020
- Technical documentation Sep 18, 2020
- A new version of nnetsauce, and a new Techtonique website Sep 11, 2020
- Back next week, and a few announcements Sep 4, 2020
- Explainable 'AI' using Gradient Boosted randomized networks Pt2 (the Lasso) Jul 31, 2020
- LSBoost: Explainable 'AI' using Gradient Boosted randomized networks (with examples in R and Python) Jul 24, 2020
- nnetsauce version 0.5.0, randomized neural networks on GPU Jul 17, 2020
- Maximizing your tip as a waiter (Part 2) Jul 10, 2020
- New version of mlsauce, with Gradient Boosted randomized networks and stump decision trees Jul 3, 2020
- Announcements Jun 26, 2020
- Parallel AdaOpt classification Jun 19, 2020
- Comments section and other news Jun 12, 2020
- Maximizing your tip as a waiter Jun 5, 2020
- AdaOpt classification on MNIST handwritten digits (without preprocessing) May 29, 2020
- AdaOpt (a probabilistic classifier based on a mix of multivariable optimization and nearest neighbors) for R May 22, 2020
- AdaOpt May 15, 2020
- Custom errors for cross-validation using crossval::crossval_ml May 8, 2020
- Documentation+Pypi for the `teller`, a model-agnostic tool for Machine Learning explainability May 1, 2020
- Encoding your categorical variables based on the response variable and correlations Apr 24, 2020
- Linear model, xgboost and randomForest cross-validation using crossval::crossval_ml Apr 17, 2020
- Grid search cross-validation using crossval Apr 10, 2020
- Documentation for the querier, a query language for Data Frames Apr 3, 2020
- Time series cross-validation using crossval Mar 27, 2020
- On model specification, identification, degrees of freedom and regularization Mar 20, 2020
- Import data into the querier (now on Pypi), a query language for Data Frames Mar 13, 2020
- R notebooks for nnetsauce Mar 6, 2020
- Version 0.4.0 of nnetsauce, with fruits and breast cancer classification Feb 28, 2020
- Create a specific feed in your Jekyll blog Feb 21, 2020
- Git/Github for contributing to package development Feb 14, 2020
- Feedback forms for contributing Feb 7, 2020
- nnetsauce for R Jan 31, 2020
- A new version of nnetsauce (v0.3.1) Jan 24, 2020
- ESGtoolkit, a tool for Monte Carlo simulation (v0.2.0) Jan 17, 2020
- Search bar, new year 2020 Jan 10, 2020
- 2019 Recap, the nnetsauce, the teller and the querier Dec 20, 2019
- Understanding model interactions with the `teller` Dec 13, 2019
- Using the `teller` on a classifier Dec 6, 2019
- Benchmarking the querier's verbs Nov 29, 2019
- Composing the querier's verbs for data wrangling Nov 22, 2019
- Comparing and explaining model predictions with the teller Nov 15, 2019
- Tests for the significance of marginal effects in the teller Nov 8, 2019
- Introducing the teller Nov 1, 2019
- Introducing the querier Oct 25, 2019
- Prediction intervals for nnetsauce models Oct 18, 2019
- Using R in Python for statistical learning/data science Oct 11, 2019
- Model calibration with `crossval` Oct 4, 2019
- Bagging in the nnetsauce Sep 25, 2019
- Adaboost learning with nnetsauce Sep 18, 2019
- Change in blog's presentation Sep 4, 2019
- nnetsauce on Pypi Jun 5, 2019
- More nnetsauce (examples of use) May 9, 2019
- nnetsauce Mar 13, 2019
- crossval Mar 13, 2019
- test Mar 10, 2019

Comments powered by Talkyard.