Modified Delphi Consensus
Progress to Date
A summary of the consensus process so far: what each round asked, what the expert panel recommended, and what remains before the framework is finalized.
Timeline
-
November 2025 — March 2026
Round 1 — Open elicitation Complete
78 of 134 invited experts completed Round 1. Three purposes for data collection were endorsed: Guiding Therapeutic Practices (Clinical), Academic Research & Knowledge Development (Academic), and Economics & Health Policy (Economic). “By Timing” was endorsed as the preferred timepoint structure. ~26,000 words of free-text suggestions were coded into a hierarchical item library.
-
April 2026 — June 2026
Round 2 — Guideline harmonization Complete
63 of 78 experts (82.1%) evaluated six structured proposals for classifying items aligned with 12 established guidelines. Round 2 consensus patterns show strong support for designating widely accepted guideline-based items as Core, a clear preference for purpose–specific approaches (Equity, Group, Economics) over universal inclusion, and an unclear deliberation regarding ReSPCT (experts agreed on the need for alignment but did not reach consensus on how); to be re-evaluated in Round 3.
-
Underway
Round 3 — Item-level evaluation In Progress
Structured ratings of the remaining candidate items that do not yet have a Core verdict, differentiated by purpose and timepoint.
-
Upcoming
Round 4 — Final consensus Planned
Final thresholds, Core vs. Supplementary verdicts, and ratification of the StaMPS Data Framework and interactive tool.
Round 1
Round 1 — What the panel established
An international, multidisciplinary panel
Of 134 invited experts, 78 completed Round 1. The panel spanned 13 countries of residence and included predominantly doctorate-level professionals: 86% reported research expertise, 59% clinical expertise, and 22% policy or health-economics expertise. Participants were most familiar with psilocybin (94%), MDMA (72%), and ketamine (71%), and collectively reported having influenced data-collection decisions for an estimated ~11,600–32,300 individuals.
Three purposes of data collection, strongly endorsed
Experts rated the appropriateness of anchoring the framework to three purposes — Guiding Therapeutic Practices (Clinical), Academic Research & Knowledge Development (Academic), and Economics & Health Policy (Economic). The mean rating was 6.52/7, with ~95% of responses in the two most favourable categories — clear consensus to proceed.
A common definition of timepoints
From four candidate categorizations of data-collection timepoints, the panel selected "By Timing" — Outset, Baseline, During Treatment, and Post-Treatment — which received both the highest mean usefulness rating (5.96/7) and the most first-choice selections (~44%). The closely rated "By Treatment Phase" option (~40%) is used selectively where it adds conceptual nuance.
Key factors shaping data collection
From 67 experts' open-text responses, 23 unique factors were identified and organized into four thematic categories (participant-level factors, intervention design and delivery, research design and outcome assessment, and organizational/contextual implementation). The most frequently cited factors were operational feasibility & resource availability (62.7%), participant burden (56.7%), and study objectives & trial phase (44.8%) — suggesting data strategies are driven above all by what teams can reliably execute and what participants can reasonably tolerate.
From 26,000 words to a structured item library
Experts proposed Core and Supplementary data domains in free text — roughly 26,000 words in total. Two independent raters coded the full dataset line by line, yielding 1,233 candidate domains that were consolidated through an iterative consensus process into 393 unique items, organized hierarchically into categories and subcategories with up to two levels of sub-items. Of these, 126 items were endorsed as potential Core measures by more than 10% of experts, and 20 by more than half. Because items were proposed spontaneously from memory, even modest counts represent meaningful independent convergence.
Mapping to existing guidelines
All items were then coded against 12 potentially relevant guidelines identified from participant suggestions and a literature review. Overlap was greatest with Good Clinical Practice (GCP; 23.7% of items) and Consolidated Standards of Reporting Trials 2025 and its extensions (CONSORT 2025; 22.1%), with partial overlap across U.S. Food and Drug Administration (FDA) Psychedelic Research Guidance (2023 draft), Reporting Setting in Psychedelic Clinical Trials (ReSPCT), CONSORT-Equity, Template for Intervention Description and Replication (TIDieR) and the Workgroup for Intervention Development and Evaluation Research (WIDER), Group-Based Behaviour-Change Interventions Checklist (GB-BCI), Consolidated Health Economic Evaluation Reporting Standards (CHEERS), and the Criteria for Reporting and Evaluating Ecological and Dynamic Intervention Components 2 (CReDECI 2) — setting up Round 2.
Round 2
Round 2 — Harmonizing with established guidelines
Round 2 asked whether items aligned with established guidelines could be readily classified as Core, using six structured proposals (599 guideline matches across the item library). It served as an evidence-informed filtering step ahead of more nuanced item-level decisions in later rounds. 63 of 78 experts completed the round (82.1%).
| Proposal | Question | Outcome |
|---|---|---|
| 1. CONSORT 2025 & extensions | Tentative items corresponding to CONSORT, CONSORT-NPT, and CONSORT-SPI guidelines should be classified as Core across all data collection purposes. | Accepted Very strong agreement. CONSORT-Equity was more mixed, with slight preference for an optional "Equity" module over universal Core status. |
| 2. GCP & FDA guidance | Tentative items corresponding to GCP recommendations and FDA Psychedelic Research Guidance should be classified as Core across all data collection purposes. | Accepted Strong agreement for GCP; moderate-to-strong for FDA guidance, with comments favouring selective, context-sensitive use. |
| 3. TIDieR / WIDER | Tentative items corresponding to TIDieR or WIDER guidelines should be classified as Core across all data collection purposes, and to fully align with TIDieR, two new Core items should be added for personalizing/titrating/adapting interventions and modifying interventions. | Accepted Strong agreement; two new Core items added. |
| 4. Set & setting (ReSPCT) | Five mutually exclusive options for harmonizing ReSPCT with the current study: (1) do not use ReSPCT to modify preliminary items; (2) classify all ReSPCT-aligned items (physical-setting and subjective-experience) as Core; (3) classify only physical-setting ReSPCT-aligned items as Core; (4) integrate ReSPCT by adding missing items and rewording for alignment; or (5) separate ReSPCT by removing overlaps and recommending its completion instead. | Re–evaluate in Round 3 ReSPCT was rated as being most relevant for clinical and academic purposes. Experts agreed on the importance of aligning with ReSPCT guidelines; however, no clear consensus emerged on how best to achieve this. This will be revisited in Round 3. |
| 5. Health economics (CReDECI 2 & CHEERS) | Tentative items corresponding to CReDECI 2 or CHEERS guidelines should be classified as Core for Economic data collection purposes only. | Economic accepted Very strong agreement — endorsement for Economic purposes only. |
| 6. Group interventions (GB-BCI) | GB-BCI–aligned items should be Core for at least one purpose, but only in group-treatment models, or Core only when therapy is delivered in groups (via a selectable "Group" option)? | Group-only accepted Universal inclusion was generally rejected; the conditional "Group" option received strong agreement and is implemented in the interactive tool. |
Consensus patterns
Three patterns emerged across proposals: strong support for making widely accepted guideline-based items Core; a clear preference for conditional or modular approaches (Equity, Group, Economics) over universal mandates. Recurring free-text themes included concern about respondent burden, the value of flexibility across study types and purposes, and support for selective rather than blanket adoption of external frameworks.
Your view
Comment on the process
We welcome observations on the process to date — including the proposals, thresholds, and how results are summarized here.