Lind et al Baseline Methods (2023)
Source details
- Type
- Paper
- Publisher
- Elsevier (Utilities Policy)
- Author
- Leandro Lind, José P. Chaves, Orlando Valarezo, Anibal Sanjab, Luis Olmos
- Published
- 2023
“Baseline methods in the context of modern distributed flexibility: an evaluation considering multi-DER types, markets, and product characteristics” — Lind et al. (2023), Utilities Policy (DOI jup.2023.101688; the raw is a preprint and names no journal — the journal is inferred). A systematic evaluation of ten baseline methodologies for DER flexibility market participation, with a proposed decision framework for method selection.
Institutions: IIT-ICAI School of Engineering, Universidad Pontificia Comillas (Madrid); VITO/EnergyVille (Belgium)
DOI: 10.1016/j.jup.2023.101688
Funding: CoordiNet project (Horizon 2020, No. 824414); BeFlex project (No. 101075438)
Summary
The problem addressed: when a DER is activated for explicit flexibility, how do you determine how much flexibility was actually delivered? For large scheduled generators, the answer is easy — compare metered output to the committed schedule. For DERs, no individual schedule exists, so a counterfactual “what would this resource have done without activation?” must be estimated. This counterfactual is the baseline.
The paper evaluates ten established methods (Table 1 lists ten rows; the raw contains no “nine” — an earlier version invented a “nine” header and a source miscount) against three criteria (accuracy, simplicity, integrity) across four DER types (Load-DR, Controllable DG, Non-controllable DG, Energy Storage Systems), multi-DER aggregation, and two product dimensions (direction: up/down; timing: real-time to weeks ahead). The central finding is that no one-size-fits-all baseline method exists.
Ten baseline methods
| Method | How it works | Best for | Key weakness |
|---|---|---|---|
| XofY | Average of X highest/mid/lowest days from the last Y eligible days | Load-DR (upward) | Upward bias (HighXofY); fails for weather-dependent DG/ESS |
| Rolling average | Average of last X same-type days (weekday/weekend), recency-weighted | Load-DR | Same as XofY; doesn’t capture DG/ESS variability |
| Comparable day | FSP selects an ex-post non-activation reference day | Non-controllable DG | High integrity for non-controllable DG (Low for Load-DR, controllable DG and ESS); FSP chooses reference day |
| Regression | Statistical model (consumption = f(weather, season, past data)) | Load-DR, PV/wind with weather data | Complex; high simplicity cost |
| Machine learning | Neural network / ML techniques | Non-controllable DG, Load-DR | Very low simplicity; black-box risk |
| MBMA | Meter reading immediately before activation = baseline | Balancing services (short-duration) | Integrity risk for ESS (see below); inaccurate for long activations |
| Zero baseline | Baseline = 0; all production during activation = flexibility delivered | Backup generators, batteries providing upward production flexibility | Fails for consumption-side DR |
| Control group | Average of similar non-activating customers during activation | Multi-DER aggregation | Requires a valid comparison group; low integrity |
| Capacity limitation | Product defined as a power cap; no energy-delta baseline needed | DSO congestion management | Requires different clearing algorithms; primarily upward only |
| Self-reported | FSP reports its own baseline | Large industrial FSPs | Low integrity without verification |
Key analytical findings
MBMA integrity risk for batteries
MBMA (Meter-Before-Meter-After) reads the meter immediately before activation and uses that reading as the baseline. For batteries, this creates a manipulation opportunity. Per raw §3.1, storage owners may have the incentive to momentarily change the battery’s state before activation only to modify the baseline (e.g. change from discharging to charging before an upward activation). (An earlier version had the direction reversed and claimed a “negative/zero baseline”, which the raw does not say.)
Zero baseline is “an alternative” for batteries in general (Table 2 rates it Medium/High/Medium for accuracy/simplicity/integrity vs MBMA Medium/High/Very Low); the raw makes no consumption/production-side split — wiki reading: — any injection during activation counts as delivered flexibility, with no pre-activation baseline manipulation possible. This is indeed the approach used in SWITCH for batteries providing increased production. (Source - SWITCH User Documentation (2026))
Capacity limitation products and the baseline question
Capacity limitation products (where the DSO sets a power cap and the FSP must stay below it) appear to eliminate the baseline problem: the product is defined by the cap, not by an energy delta. Per raw §2, a capacity limitation product eliminates the need for a baseline (Table 1 lists it as static, with no defined baseline data); the raw’s only qualifiers are that clearing algorithms must differ and that it is primarily for upward congestion relief. (An earlier version claimed the authors note such products “often still require energy delivery validation” — not in the paper; that was a wiki extrapolation from SWITCH.)
Harmonisation across sequential markets
When a DSO LFM and a TSO balancing market operate sequentially (both drawing from the same portfolio), different baseline methods create distortions. (The TSO-MBMA vs DSO-XofY conflicting-incentives scenario is a wiki extrapolation; the raw says only that harmonisation across sequential TSO/DSO markets avoids distortions and gives the FSP certainty on remuneration.) The paper recommends baseline harmonisation across interacting markets — relevant to future TSO-DSO coordination as NC DR matures. (Network Code on Demand Response)
Market timing
- Real-time / balancing services: MBMA is “the most used baseline method for balancing services” (DNV-GL 2020; the raw does not link it to FCR/mFRR specifically); no time for ex-ante calculation.
- Day-ahead cleared products: XofY or rolling average; excluding the hours between GCT and activation is only “an alternative” — those hours may still matter for accuracy, so procurers need rules to verify data (e.g. Elia: no adjustment by default; on request a three-month evaluation; justification if the adjustment raises the baseline by more than 15%).
- Long-term contracted products (ST/LongFlex): ex-ante calculation methods; regression or ML feasible.
DER-type matrix (summary)
- Load-DR: historical methods (XofY, rolling average) are adequate — medium accuracy, high simplicity.
- Non-controllable DG (wind, solar): regression or ML needed for accuracy (weather-driven output); XofY only works with same-day adjustment.
- Controllable DG (backup generators): zero baseline is the most accurate (baseline IS zero when idle).
- ESS: MBMA is technically accurate but its integrity is at great risk; zero baseline is an alternative (no consumption/production split in the raw).
- Multi-DER aggregation: no single method covers mixed portfolios well; submetering per technology type is the most accurate but costly; comparable day or control group are pragmatic alternatives.
Connections to Swedish context
- SWITCH MBMA — the default automatic baseline in SWITCH for consumption-side resources. The paper confirms this is the correct approach for short-duration balancing-type products, though the raw says only “low accuracy in long activations, given that the baseline cannot be changed” (no 1–2 hour threshold). (Wiki inference: the paper never mentions SWITCH; it does not “confirm” or endorse SWITCH’s design.)
- SWITCH zero baseline (noll-referens) — used for battery resources providing production-side upward flexibility. (Wiki inference: the paper never mentions SWITCH; it rates zero baseline for batteries Medium/High/Medium.)
- NODES rolling average — sthlmflex used a 5-day rolling average as the standard NODES baseline (Source - sthlmflex säsong 3 (2022-2023)). The paper categorises this as a rolling average variant with medium accuracy and medium integrity for Load-DR.
- NC DR (article numbers not in the paper) settlement — future flexibility markets under NC DR will require standardised baseline methods. (Wiki inference; the raw cites no NC DR article, only the ACER Framework Guideline on Demand Response.)
- BeFlexible project — the raw names the funder as “BeFlex Project No. 101075438”; the link to SWITCH demonstrations is outside knowledge. The academic framework and the operational SWITCH design are therefore closely related.
Relevance to wiki topics
| Topic | Relevance |
|---|---|
| Baseline Methods | Primary source for the concept page |
| SWITCH | MBMA and zero-baseline methods explained; capacity/energy distinction clarified |
| NODES | NODES uses rolling average in sthlmflex; no MBMA |
| Aggregation | Multi-DER baseline challenge is the core aggregation settlement problem |
| Energy Storage | Battery-specific baseline recommendations; MBMA integrity risk |
| Flexibility Market | Baseline design is central to market settlement and FSP participation |
| Network Code on Demand Response | Future regulation will standardise baseline methods; NC DR (article numbers not in the paper) |
| CoordiNet | Paper funded by CoordiNet; SWITCH’s MBMA baseline traces to CoordiNet design |