Skip to content

Monitoring

Squad Federation provides a hybrid monitoring model: conversational skills for quick status checks and optional OpenTelemetry dashboards for deep observability.

The federation-orchestration skill answers questions about team status, progress, and health.

“How’s my federation doing?”

The skill shows a dashboard:

📊 Squad Federation Dashboard
Team State Step Progress Updated
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
❌ backend failed analyzing routes 65% 1m ago
🔄 frontend scanning auth module 45% 2m ago
✅ infra complete - 100% 5m ago

State indicators:

  • ❌ failed - Error occurred
  • 🔄 Active states (initializing, scanning, distilling)
  • ✅ complete - Finished successfully
  • ⏸️ paused - Manually paused

“What’s the frontend team doing?”

The skill shows detailed status:

Frontend Team Status:
State: scanning
Step: analyzing authentication module
Progress: 45%
Last update: 2 minutes ago
Deliverable: ✅ Present
Recent learnings: 2

“Are any teams stuck?”

The skill warns if teams haven’t updated in >10 minutes:

⚠️ Backend team: No update in 15m (possible stall)

Check the team’s log:

Terminal window
tail -100 .worktrees/backend/run-output.log

“Show me team deliverables”

📦 Deliverables:
frontend: ✅ deliverable.md
backend: ❌ Missing
infra: ✅ OUTPUT.json

“What have my teams learned?”

📚 Recent Learnings:
frontend (3):
- pattern: Parallel test execution reduces CI time
- discovery: Auth context passed via props
- convention: Name API routes with kebab-case
backend (2):
- pattern: Use dependency injection for database clients
- gotcha: Don't import barrel files in tests

If you enabled telemetry during federation setup, you have access to the OpenTelemetry Aspire dashboard.

The dashboard runs at http://localhost:18888 (auto-started if you chose “Yes” during setup).

If it’s not running, start it:

Terminal window
docker run -p 18888:18888 -p 4318:18889 \
mcr.microsoft.com/dotnet/aspire-dashboard:latest

Traces:

  • Team lifecycle operations (onboard, launch, monitor)
  • Placement operations (read, write, commit, push)
  • Communication operations (readSignals, writeSignals)

Metrics:

  • squad.teams.active - Active team count
  • squad.teams.failed - Failed team count
  • squad.signals.sent - Signal send rate
  • squad.learnings.created - Learning creation rate

Logs:

  • Team startup/shutdown events
  • Signal send/receive events
  • Learning capture events
  • Error messages

Each team gets an .mcp.json file with telemetry config:

{
"mcpServers": {
"squad-otel": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-opentelemetry"],
"env": {
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://localhost:4318",
"OTEL_SERVICE_NAME": "squad-frontend"
}
}
}
}

Teams emit telemetry via the OpenTelemetry MCP server during their sessions.

Teams maintain .squad/status.json with current state:

{
"state": "scanning",
"step": "analyzing authentication module",
"updated_at": "2025-01-30T12:30:00Z",
"agent_active": "lead",
"progress_pct": 45,
"error": null
}

States:

  • initializing - Team starting up
  • scanning - Analyzing codebase
  • distilling - Processing findings
  • complete - Finished successfully
  • failed - Error occurred
  • paused - Manually paused

Ask the orchestration skill to show detailed status for a specific team:

“What’s the frontend team status?”

The status file is also available at .worktrees/frontend/.squad/signals/status.json.

Each archetype can provide custom monitoring logic beyond the basic status check.

Archetype plugins include monitors that emit domain-specific metrics:

Coding archetype:

  • Files changed count
  • Test count
  • PR readiness checks

Deliverable archetype:

  • Deliverable completeness
  • Schema validation status
  • Output file size

Consultant archetype:

  • Questions answered
  • Domains indexed
  • Insights provided

These metrics flow to the telemetry dashboard if enabled.

When teamsConfig is set in federate.config.json, the meta-squad posts status summaries to a Microsoft Teams channel automatically. This works through the skill layer — Copilot sessions have native access to the Teams MCP tools. The teams-presence feature provides automated, persistent monitoring by running as a background bridge process that polls the channel and relays messages to a Copilot ACP session.

  • Status summaries — periodic status updates for all teams via teams-presence
  • Directive relays — confirmation when directives are sent to teams
  • Alert notifications — team failures, stalls, or critical errors

You can post messages tagged with @<federationName> in the configured Teams channel. The teams-presence feature polls for these and acts on them:

@<federationName> tell frontend to skip legacy utils
@<federationName> pause backend
@<federationName> restart infra

Add teamsConfig to federate.config.json:

{
"teamsConfig": {
"teamId": "your-teams-team-guid",
"channelId": "19:your-channel-id@thread.tacv2"
}
}

The meta-squad uses the PostChannelMessage MCP tool to post and ListChannelMessages to poll. No additional setup is needed — these tools are available natively in Copilot sessions.

For full details on the persistent bridge, see the Teams Presence guide.

  • You want status updates without keeping a terminal open
  • Multiple stakeholders need visibility into federation progress
  • You want to send directives from Teams instead of the terminal

For setup details, see the configuration reference.

Ask the orchestration skill:

“Why isn’t the frontend team showing?”

The skill will check the team registry and help diagnose the issue.

If the team is missing, re-run onboarding:

“Onboard a team for frontend”

Ask the orchestration skill to check the team:

“Why is the frontend team stalled?”

The skill will check the session status and error logs.

To restart:

“Restart the frontend team”

Ask the orchestration skill to check telemetry:

“Is telemetry working for my teams?”

The skill will verify the dashboard is running and teams are exporting telemetry.

If telemetry wasn’t enabled, you can enable it by running federation setup again and choosing telemetry, then restarting teams.