The decision to go multi-site is usually made too late. A single-site study finishes, the manuscript goes out, and a reviewer writes "small sample, single centre, limited generalisability". Adding a second site at that point does not rescue the study: different selection, different measurement, different time window, two datasets that will not merge. The moment to design a multi-site study is before the first participant.
The hard part is administrative, not scientific
Thinking of a multi-site study as "more participants" is the most common error. The scientific question stays the same; what changes is governance. The difficulty concentrates in four places: making the protocol precise enough for another team to run, obtaining ethics approval at every site, collecting data against one shared dictionary, and agreeing authorship in writing before the work starts. None of these requires scientific skill. All of them require time, and all of them are easy to postpone.
Choosing sites: complementarity, not prestige
The first question about a second site is not how large its name is but what it brings that you do not have. Useful complementarities:
- A different participant profile. Different region, different socioeconomic distribution, different referral pattern. This is where the generalisability claim comes from.
- Case volume. Ask for the site's actual number from last year, taken from its records, not an estimate given in a meeting.
- Capability you lack. Equipment, a laboratory, a measurement method, or an expertise that exists there.
- A named person who will do the work. Commitment comes from people, not institutions. Every site needs a named local lead.
Few and solid usually beats many and loose. A two-site study with clean data is easier to publish than a six-site study where three sites never send anything.
A protocol that can be run without asking you
The test is simple: if you handed the protocol to a team you have never spoken to, could they run the same study without asking a question? If not, the protocol is incomplete. What has to be unambiguous:
- Inclusion and exclusion criteria, with their cut-off values
- The primary outcome and exactly how it is measured: which instrument, by whom, when
- Whether assessors are blinded, and how
- What happens with missing data and protocol deviations
- How measurement drift between sites will be monitored
Ready-made frameworks make this easier. For clinical trial protocols, SPIRIT is written for protocol authoring itself[1]. At the reporting stage, CONSORT covers randomised trials[2] and STROBE covers observational studies[3]. Those two are reporting guidelines, not study design or quality appraisal tools; even so, opening the checklist while writing the protocol rather than the manuscript exposes gaps you could not fix later.
Multi-site work outside the clinic
The same logic holds when there are no patients; only the vocabulary changes. In an inter-laboratory measurement study, "site difference" means instrument and calibration difference, and the shared data dictionary is replaced by a shared measurement protocol and reference samples. In field research the variability comes from the data collectors: training and a pilot are what stop the same question being asked in three different ways. In multi-institution data projects the critical heading is data management, with variable names, units and versioning fixed at the start. The FAIR principles give a shared vocabulary for that[4]. Ethics requirements vary by field, but where there are human participants or personal data, institutional review is generally required outside clinical research too.
Ethics approval and data sharing
The usual sequence: the coordinating site obtains approval, the protocol and that approval travel to the other sites, and each site obtains approval from its own committee or IRB. This varies between institutions and countries and never takes less time than you expect; put several months in the timeline.
Data sharing is a separate heading and cannot be left implicit. No field containing personal data should move between sites in raw form: data is de-identified at its own site and enters the shared pool with a code only. Who can access what, where data is stored, and what happens to it after the study must be written down. Cross-border transfer brings its own legal requirements; in the European Union the GDPR governs this[7], and other jurisdictions have their own equivalents.
Authorship: in writing, before the first participant
Multi-site studies most often crack over author order rather than results. What should be written down before the study starts: what contributions are expected to qualify for authorship, which principles determine the order, how contributions will be recorded during the study, and how a site joining later fits into those principles. The author list itself is settled at the end on the basis of contributions that actually happened; allocating a quota per site in advance conflicts with that. Naming contributions with CRediT roles makes the conversation concrete[5], and in biomedical fields the ICMJE criteria state that data collection alone does not qualify someone for authorship[6]. Saying this at the start is easier than saying it at the end.
Red flags
- "We will sort out the data format later." Data collected without a shared dictionary does not merge later; it does not merge at all.
- Enthusiastic verbal commitment with no named local lead.
- Case volume that was never verified, only estimated in a meeting.
- Assuming the coordinating site's approval is enough. In most institutions it is not.
- Starting without a pilot. Five cases will expose the questions your data form asks ambiguously, before that ambiguity spreads across the whole study.
What must be finished before the first participant
- Protocol read by every site, with their questions closed
- A named local lead at each site
- One shared data dictionary and a single data collection form
- A five-case pilot and the form revision that follows it
- Coordinating site approval obtained, other applications submitted
- Data sharing and retention rules written down
- Authorship principles written down and accepted by all sites
- Sample size calculation divided into a per-site target
Sources and standards
- SPIRIT — standard protocol items for clinical trials. equator-network.org
- CONSORT — reporting of randomised controlled trials. equator-network.org
- STROBE — reporting of observational studies (a reporting guideline, not a design or quality appraisal tool). strobe-statement.org
- FAIR principles — findable, accessible, interoperable, reusable data. go-fair.org
- CRediT (Contributor Roles Taxonomy), NISO. credit.niso.org
- ICMJE, "Defining the Role of Authors and Contributors". icmje.org
- GDPR — personal data and international transfers in the EU. gdpr.eu
The hardest step is finding the right second site
KONSİL derives the roles a research idea requires and lists candidate researchers for each role from open publication data, showing their in-topic output, institution and shared language. If the design calls for multiple sites, the report says so. The sample report shows how five roles and their candidates were produced.
Apply for the closed beta