Canonical definition

Collective-State Inference (CSI) is the process by which an AI system infers latent, emergent conditions of a bounded collective—such as engagement, cohesion, conflict, and alignment—from observable interaction patterns among multiple participants within a shared environment. These conditions arise, stabilize, or change through collective dynamics and are defined at the level of the collective rather than assumed to be reducible to aggregations of individual states, attributes, or behaviors.

CSI treats the collective itself—not merely its individual members, an external environment, or a descriptive group metric—as the object of inference.

Claim status: The definition and qualification requirements are original propositions of the CSI research program. They synthesize established multilevel, emergence, composition, and construct-validity scholarship but have not yet completed peer review or empirical validation.

Why the definition matters

A computational output needs an argument connecting observations to the claimed collective condition. The same interaction pattern may have different interpretations. CSI specifies the required collective boundary, target construct, composition logic, temporal and contextual alignment, uncertainty, and collective-level validation.

CSI's proposed contribution: CSI connects the collective boundary, target construct, composition logic, evidence, time, context, uncertainty, and validation into an account of when computational outputs warrant collective-state interpretations. This is a theoretical proposal; empirical validation remains to be completed.

Reasoning and proposed evaluation

Established foundations and proposed contribution: Published work already models dynamic group affect (Prabhu et al., 2025) and team constructs through temporal and relational representations (De Luca et al., 2026). CSI’s proposed contribution is to specify, across construct families, when such outputs warrant interpretation as latent conditions of bounded collectives: explicit boundaries, construct level, composition, history, context, uncertainty, and validation. These approaches are close comparators. Evaluating CSI therefore requires evidence of construct validity, calibration, and incremental value where theory predicts it, rather than predictive accuracy alone.

Predicting a group outcome does not by itself establish that a model estimates its collective condition. CSI therefore proposes comparisons among aggregate baselines, relational models, and fuller CSI models using independent collective-level criteria.

CSI predicts added validity from appropriate composition and relational evidence when the target construct requires that structure. Failure to improve on simpler models can inform the construct definition, evidence, time window, or theoretical account.

What CSI means

CSI addresses a specific inferential problem: how an AI system can estimate a condition that exists at the level of a collective but is not directly observable. The system therefore reasons from evidence—such as communication, coordination, reciprocity, subgroup structure, persistence, and context—to a latent collective-level representation.

The inferred state is not assumed to be a hidden fact that can be read directly from data. It is a theory-guided, probabilistic estimate whose meaning depends on how the collective is bounded, how the construct is defined, how lower-level evidence composes into a collective-level condition, and how the resulting estimate is validated. The separation of lower-level observations from higher-level constructs and the requirement for an explicit composition model are established in multilevel theory (Chan, 1998; Morgeson & Hofmann, 1999; Kozlowski & Klein, 2000).

Contemporary computational neighbors and the CSI novelty boundary

Contemporary research already includes explicit computational targets at the group or team level. Prabhu et al. (2025) model dynamic group affect from group-level annotations and multimodal interpersonal synchrony; De Luca et al. (2026) jointly model temporal interactions and relations while predicting team constructs; Proutskova (2026) applies collective-state inference to reciprocal coordination in human–AI vocal ensembles; and Riedl (2026) studies emergent coordination in language-model agent systems. These works make it inappropriate to claim that CSI is the first computational approach to represent or predict a group or collective state.

Neighboring focusQuestion CSI carries across construct families
Dynamic group affect and interpersonal synchronyWhen does multimodal group evidence warrant interpretation as a latent collective condition?
Temporal-relational team modelingWhich composition, comparison, and validation conditions justify the team-level interpretation?
Domain-specific collective-state inferenceWhich inferential commitments transfer across settings, constructs, and collective types?
Emergent coordination among AI agentsHow should collective boundaries, uncertainty, and collective-level validity be specified?

CSI's proposed contribution is theoretical: it treats collective-state estimation as a general, level-aware class of AI inference and specifies the conditions under which heterogeneous computational outputs may legitimately be interpreted as latent conditions of bounded collectives. The framework therefore asks whether the target exists at the collective level, whether the composition rule matches the construct, whether evidence is aligned to level and time, whether uncertainty is represented, and whether construct, discriminant, incremental, and boundary validity are established. Recent group-state models are close comparators. Their existing construct and validation commitments must be examined before claiming that CSI adds an inferential requirement or consequence they do not already provide.

When an inference qualifies as CSI

An AI-generated group assessment does not qualify as CSI merely because it summarizes several participants or directly predicts a group-level variable. Four requirements distinguish CSI from ordinary aggregation, reporting, task-specific group prediction, or individual-level prediction. CSI proposes these requirements as an integrated qualification rule derived from the multilevel distinction between constructs, measures, and composition rules (Chan, 1998; Morgeson & Hofmann, 1999).

1. Explicit boundary

Bound the collective

The participants, shared environment, relevant context, and observation window must be specified.

2. Collective-level construct

Define the target condition

The target must be a property of the collective rather than a relabeled individual, dyadic, or organizational measure.

3. Justified composition

Explain how evidence composes

The estimate must use a theoretically appropriate composition model rather than unexamined aggregation.

4. Collective-level validation

Test the estimate

The inferred state must be compared with collective-level criteria and simpler aggregate or context-free baselines.

The CSI inferential chain

1. Bound the collective

CSI begins by defining who belongs to the focal collective, the environment they share, the relevant task or context, and the time interval over which the condition is being estimated.

2. Specify the collective state

The target construct must be defined at the collective level, with an appropriate timescale and composition model. Engagement, cohesion, conflict, and alignment are examples, but each requires its own theoretical specification.

3. Integrate theory-aligned evidence

The system uses evidence justified by the construct. Depending on the state, this may include participant-level traces, relational patterns, temporal dynamics, and contextual information.

4. Estimate state and uncertainty

The output should represent the most plausible collective condition, credible alternatives, and uncertainty arising from boundaries, evidence, models, and temporal change.

5. Compare and validate

CSI proposes triangulated validation using group-referenced surveys, trained observer ratings, independently coded interaction episodes, and comparisons with simpler models. These are proposed validation requirements, not completed empirical results.

6. Support collective awareness

Validated CSI outputs may support systems that interpret, project, explain, and reason about collective conditions under appropriate human oversight and governance.

What CSI is not

Adjacent approachWhat it doesWhy it is not automatically CSI
Individual analyticsEstimates the state, preference, intent, or behavior of individual participants.The individual rather than the collective is the object of inference.
AggregationCombines individual scores, counts, or attributes into a group statistic.A statistic does not establish a latent collective construct without justified composition and validation.
Group analyticsDescribes participation, volume, sentiment, network measures, or performance.Descriptive metrics may be useful evidence but need not represent an inferred collective state.
Group-state or team-state modelingDirectly predicts group states, team constructs, or collective dynamical regimes from pooled, relational, temporal, or multimodal evidence.A group-level target may be compatible with CSI, but qualification still depends on construct definition, composition logic, uncertainty, collective-level validation, and stated boundaries.
Relational collective inferenceJointly predicts labels or states of interconnected entities.The targets remain entity-level labels rather than an emergent condition of the bounded collective.
Distributed state estimationMultiple agents jointly infer an external or environmental state.The collective performs the inference; it is not necessarily the object being inferred.
AI summarizationProduces a narrative description of group interactions.A fluent summary does not demonstrate construct definition, composition logic, uncertainty, or validation.

Relationship to collective awareness

CSI is the inferential mechanism. Collective awareness is the broader functional capacity that may be built on calibrated CSI outputs. A collective-aware AI system may interpret what a state means, explain the evidence supporting it, project how it may change, and help people reason about possible responses. This functional framing is informed by situation-awareness research while remaining a distinct CSI construct (Endsley, 1995).

CSI therefore does not imply consciousness, sentience, or autonomous authority. It provides a disciplined representation of a collective condition that can support human judgment when uncertainty, limitations, and governance requirements remain visible.

How this article connects

RA-001 is the conceptual gateway into the CSI research program.

The relationships below show how the canonical CSI definition connects to its theoretical foundation, core concepts, measurement requirements, functional extension, and research instrumentation.

These links represent conceptual and programmatic relationships. They do not imply formal citation direction, peer-review status, or equivalent evidentiary maturity.

References and scholarly foundations

The following works support the multilevel, construct-validity, and contemporary-neighbor boundaries used in this article. CSI’s proposed contribution is the integration of four qualification requirements, a six-stage inferential chain, and a cross-construct specification. The component ideas draw on the cited literature; their integration and claimed consequences remain to be tested.

Chan (1998) — composition models connecting lower-level observations to higher-level constructs.

Morgeson and Hofmann (1999) — the structure and function of collective constructs.

Kozlowski and Klein (2000) — contextual, temporal, and emergent multilevel processes.

Sawyer (2004) and Sawyer (2005) — mechanisms and social-system accounts of emergence.

These validation requirements are CSI research-program proposals.

Prabhu et al. (2025), De Luca et al. (2026), Proutskova (2026), and Riedl (2026) — close computational comparisons that bound CSI’s proposed contribution.

Complete bibliographic details and stable identifiers are available in the Scholarly Sources registry. Citation practices are governed by the Editorial and Citation Policy.

Research status

RA-001 documents the current CSI research-program position. It remains a foundational draft subject to scholarly review, empirical operationalization, and validation. The proposed capabilities and expected benefits have not been empirically established.

Preferred interim citation

Clark, E. D. (2026). Collective-State Inference. CSI Reference, RA-001, Version 0.9. Collective-State Inference Research Program.

See the Reference version history for release status and substantive changes.