# FireFight > The open-source incident management platform built entirely for Slack. Fast, transparent, and developer-first. This file is the full text of every page on firefight.app, concatenated for AI clients. Each section is preceded by the canonical URL of the page it came from, which is the URL to cite. A linked index of the same pages is at https://firefight.app/llms.txt. --- Source: https://firefight.app/ # FireFight **Code breaks. Fix it without leaving Slack.** Open-source incident management for Slack. One command opens the channel, pulls in the right people, and captures everything as it happens. When it is over, the postmortem is already drafted. - Open source - Built for Slack - Agent ready [Sign in with Slack](https://app.firefight.app/) 14 days free, no credit card. Viewers never cost anything. ## Why teams pick FireFight ### AGPL-3.0, the whole thing Not an open core with the useful parts held back. Read it, change it, keep the changes. ### Your servers or ours Let us run it and never patch or upgrade a thing. Or self-host the identical platform if you would rather. ### One flat price, not per-seat math Plans are sized by responders, the people who act on incidents. Everyone else follows along in Slack for free. ### Integrates with anything A versioned REST API, outbound webhooks, and an MCP server your AI agents can act through. ## Incident management shouldn't be this painful Your team already lives in Slack. Incidents come with enough pressure. The process should not add more. ### A stack held together with glue An alerting tool, a chat thread, a doc template, a status page. You stitch them together with a process doc, and it comes apart in the heat of the moment. ### The first 20 minutes go to setup Opening the channel, finding the right people, posting the first update. By the time everyone is in the room, the incident has had a head start. ### Answers buried in threads Critical details are trapped in DMs, threads, and dashboards. Anyone joining late has to wade through hundreds of messages to learn what everyone else already knows. ### The fix ships, the write-up waits Someone still has to rebuild the timeline, chase the follow-ups, and write the postmortem, usually after hours. So it gets skipped, and the same incident comes back two months later. ## Built to resolve. Designed for Slack. FireFight connects your tools, teams, and timelines so every incident is captured, automated, and resolved without leaving your workflow. ### Run every incident directly in Slack FireFight lives where your team already works. Create incidents, assign roles, and post updates without leaving Slack. ### Webhooks & API that keep everything in sync Subscribe to incident events with webhooks, or manage incidents programmatically through our REST API. ### Postmortems that write themselves FireFight's AI turns every incident into a clear story you can learn from, automatically pulling timelines, decisions, and Slack context into one place. ### Never wonder who owns what Define your teams, services, and environments once, and attach the runbook each one needs. When an incident starts, the owner, the context, and the steps to take are already in the channel. No asking around, no stale spreadsheet. ## Built for retries, duplicates, flapping and alert storms Monitoring tools do not send one clean alert. They send the same one five times, then four hundred at once at 3am. FireFight sorts that out before anyone is notified. ### One alert, not five However many times your monitoring resends the same alert, your team sees one and deals with one. ### Flapping stays quiet An alert that clears and comes straight back is still one problem. Nobody is notified twice for it. ### Storms become one incident When related alerts fire at once they land in one incident, with one channel and one timeline. Not forty channels with nobody in charge. ### Noise cannot bury a real incident A misconfigured monitor firing nonstop is held at the door. The alert that matters still gets through. [See how alerts are grouped](https://firefight.app/docs/alerts/grouping-and-deduplication) ## Give agents real access, without handing over the keys Agents read the same incident record a responder sees, over MCP and the REST API. Grant access capability by capability. Nothing runs unless you allowed it, and every attempt is recorded. When an answer lives in Datadog, Grafana, Sentry or your database, FireFight asks for it there. It is not another place your logs, metrics and traces have to live. Agents connect over MCP: Claude Code, Cursor, or any MCP-compatible client. ## Operational in minutes, not weeks 1. **Sign in**: Create your FireFight account. 2. **Install to Slack**: Connect FireFight to Slack. 3. **Start responding**: Declare your first incident. Average setup time: under 5 minutes. ## Don't trust us. Audit us. Every line of FireFight is open source. Read how it handles your alerts, what it lets agents do, and what it keeps. Nothing we say on this page has to be taken on faith. - Source: [github.com/FireFightLabs/firefight](https://github.com/FireFightLabs/firefight) - Hosted: [app.firefight.app](https://app.firefight.app/) - Pricing: [firefight.app/pricing](https://firefight.app/pricing) ## The things teams ask before they switch **Does everyone on the team need a paid seat?** No. Plans are flat and sized by responders, the people who act on an incident by declaring one, changing a status or severity, taking a role, or editing a postmortem. Everyone else can join the incident channel, read along, and post in the thread without ever costing anything or counting toward your plan. [See plans and pricing](https://firefight.app/pricing) **Is FireFight really open source, or is it open core?** The whole platform is AGPL-3.0. There is no separate paid edition of the source and no feature held back for the hosted version. If you self-host, you are running the same thing we run. **How long does setup actually take?** Sign in, install the Slack app, declare your first incident. There is nothing to configure beforehand, because statuses, severities, roles, and the incident form all ship with working defaults you can change whenever you want. **What does the AI do during an incident?** It keeps a running summary of the channel, so anyone joining late can run /ff catchup and be current in a few seconds. Mention it and it answers questions about the incident from that incident’s own record. When the incident closes it drafts the postmortem from the timeline and the discussion, and it will rewrite any passage of that draft when you tell it how. **Do you train AI models on our data?** No. We do not use your data to train AI models, and we do not use any provider or setting that would allow anyone else to. FireFight works one incident at a time, from that incident’s own channel and timeline, and nothing else in your workspace is touched. Known credential formats are stripped before anything is saved, and everything saved is encrypted. Self-host and none of it leaves your own infrastructure. [What is stored, what is redacted, and what reaches a model](https://firefight.app/docs/ai/data-handling) **Can AI agents change things in our systems?** Only what you have explicitly granted. Every ability is given to a named agent for a named environment, nothing is on by default, and you can take any of it back at any time. Anything high risk waits for a person to approve it. Every attempt is recorded whether it was allowed or refused, so you can always see what an agent did and what it was stopped from doing. [How permissions, approvals and the activity record work](https://firefight.app/docs/gateway/permissions) **What happens when the same alert fires ten times?** Your team sees one alert and one incident. An alert that clears and comes back is still that one alert, and a burst of related alerts becomes one incident with one channel. A monitor stuck firing nonstop is held back so a real alert still gets through. [How alerts are grouped and deduplicated](https://firefight.app/docs/alerts/grouping-and-deduplication) **Does it work with the tools we already run?** FireFight connects to GitHub, GitLab, Datadog, Grafana, New Relic, Sentry, Linear, Notion, Confluence, and Postgres on Neon, Supabase, or PlanetScale. Anything else that speaks MCP can be connected directly, and outbound webhooks plus a versioned REST API cover whatever is left. **What happens when the trial ends?** You are asked for a card. If you do not add one the workspace pauses, so no new incidents can be declared, and everything you already have stays exactly where it is until you reactivate. **What if we want to leave?** Cancel from your settings at any time. There is no contract and no cancellation fee. Your data is kept for 30 days in case you come back, then permanently deleted. If you would rather keep running FireFight without us, self-hosting is the same software. ## Your next incident is already on its way Install FireFight in your Slack today. Setup takes five minutes. [Sign in with Slack](https://app.firefight.app/). 14 days free, no credit card, cancel whenever you like. Or run it on your own servers and pay nobody at all. Not ready to install? Leave your email on the page and we will write when something ships. ## Learn more - [Blog](https://firefight.app/blog): engineering deep dives and product thinking - [Changelog](https://firefight.app/changelog): recent updates and fixes - [Pricing](https://firefight.app/pricing): flat plans, every feature on every plan --- Source: https://firefight.app/pricing # Pricing **Cheaper than one hour of downtime.** Flat plans sized by responders. Every feature on every plan, AI credits included, and viewers are always free. ## Plans Every plan includes every feature and unlimited free viewers. Scale adds priority support. Enterprise adds SAML SSO, SCIM, and self-hosting support. Every plan starts with a 14-day free trial with 100 AI credits to try the AI. Your plan's monthly allowance begins when you subscribe. A credit is 5 cents, and credits never gate a feature. ## AI credits A credit is 5 cents. Your plan includes a monthly allowance, and you top up whenever you need more. Here is what uses them. - **Catch-up.** Run /firefight catchup in an incident channel for a summary of what has been tried, decided, and blocked. - **Responder answer.** Mention @Firefight in an incident channel and get an answer grounded in the incident record. - **Postmortem draft.** A full postmortem written from the timeline and the channel, ready to edit. - **Section rewrite.** Select a passage in the postmortem editor and tell the AI how to rewrite it. - **Timeline notes.** When an incident closes, FireFight reads the channel once and writes the theories, findings, root cause, and mitigation onto the timeline. - **Resets monthly, top-ups stay.** Your allowance resets every month, and what you did not use does not carry over. Credits you buy as a top-up never expire. - **AI pauses at zero, nothing else does.** We warn you at 70% and 90%. At zero, AI features wait for the next refill or a top-up. Declaring, updating, and resolving incidents never use credits. - **Top-ups are yours to switch on.** Buy 500 credits for $25 from your workspace settings, or turn on automatic top-ups with a monthly cap you set. Nothing charges on its own. ## Compare plans | | Starter | Team | Scale | Enterprise | | --- | --- | --- | --- | --- | | Monthly price | $99 | $299 | $599 | Custom | | Responders | Up to 10 | Up to 30 | Up to 100 | Unlimited | | AI credits a month | 300 | 1,000 | 3,000 | Pooled | | Free viewers | Unlimited | Unlimited | Unlimited | Unlimited | | Priority support | - | - | Yes | Yes | | SAML SSO and SCIM | - | - | - | Yes | | Self-hosting support | - | - | - | Yes | | AI postmortem drafts | Yes | Yes | Yes | Yes | | AI Slack thread summaries | Yes | Yes | Yes | Yes | | AI similar incident search | Yes | Yes | Yes | Yes | | Runbooks | Yes | Yes | Yes | Yes | | Declare & resolve in Slack | Yes | Yes | Yes | Yes | | Incident roles | Yes | Yes | Yes | Yes | | Actions & follow-ups | Yes | Yes | Yes | Yes | | Configurable severities & statuses | Yes | Yes | Yes | Yes | | Automatic incident timeline | Yes | Yes | Yes | Yes | | Postmortem editor | Yes | Yes | Yes | Yes | | Service catalogue | Yes | Yes | Yes | Yes | | Incident metrics | Yes | Yes | Yes | Yes | | Slack channel management | Yes | Yes | Yes | Yes | | Full audit trail | Yes | Yes | Yes | Yes | | Integrations (GitHub, GitLab, Datadog, Grafana, New Relic, Sentry, Linear, Notion, Confluence, Postgres) | Yes | Yes | Yes | Yes | | MCP server for AI agents | Yes | Yes | Yes | Yes | | Agent permissions & approvals | Yes | Yes | Yes | Yes | | Webhooks & REST API | Yes | Yes | Yes | Yes | Plans differ in size and support, never in features. [Start your 14-day free trial](https://app.firefight.app/). No credit card required, cancel anytime. ## Self-host Prefer to run it yourself? FireFight is open-source. See the [self-hosting docs and GitHub repo](https://github.com/FireFightLabs/firefight). ## FAQ **What happens after the 14-day trial?** Your trial includes every feature and a one-time 100 AI credits. Need more, top up 500 for $25, and any you buy stay with you when you subscribe. At the end of the trial, you'll be asked to enter a payment method. If you don't, your workspace is paused. No incidents can be declared, but all your data is preserved. You can reactivate anytime by entering a card. **What counts as a responder?** A responder is someone who acts on an incident, not someone who watches one. You become a responder the first time you sign in to the dashboard, declare an incident, change a status or severity, post an official update, take or get assigned a role, create or complete a follow-up, or edit a postmortem. Following along in Slack stays free, forever: joining any incident channel and posting in the thread never costs anything. Responders are added automatically, so nobody is ever blocked in the middle of an incident, and you review and remove them from your dashboard. **So most of my company is free?** Yes. Viewers are unlimited on every plan. In a 100-person engineering org where 20 people run incidents, the Team plan covers you. The other 80 can follow any incident in Slack and add context in the channel without costing anything or counting toward your plan. **What are AI credits?** A credit is 5 cents. Every plan includes a monthly allowance, and AI work draws on it: catch-ups, responder answers, postmortem drafts, section rewrites, and the timeline notes written when an incident closes. The live summary FireFight keeps behind these is part of that work, not a separate charge. Your allowance resets every month and does not carry over, and top-up credits never expire. Core incident management never uses credits, so declaring, updating, and resolving incidents always works, even at zero. **What happens when we run out of credits?** You hear about it first. We warn you at 70% and 90% of your allowance. At zero, AI features pause until your next refill or a top-up, and everything else keeps working. Top up 500 credits for $25 from your workspace settings, or turn on automatic top-ups with a monthly spending cap you set. Nothing charges on its own. **What happens if we outgrow our responder limit?** Nobody is ever blocked in the middle of an incident. If your workspace grows past the responders your plan covers, we ask you to move up a plan. You can review and remove responders from your dashboard at any time. **Can I cancel anytime?** Yes. No annual contracts, no cancellation fees. Cancel from your workspace settings at any time. Data is kept for 30 days after cancellation in case you reactivate, then permanently deleted. **Do you offer annual billing?** Not yet. Monthly billing only for now. Annual billing with a discount is on the roadmap. We'll announce it in the changelog when it's ready. **Is there an enterprise plan?** Yes. If you need more than 100 responders, SAML SSO, SCIM, a pooled AI credit allowance sized to your usage, or help running FireFight in your own infrastructure, email us and we'll talk. No procurement theatre, no six-month implementations. **What integrations are included?** Webhooks and a public REST API are supported today, so you can wire FireFight into any tool that speaks HTTP. On top of that, FireFight connects directly to GitHub, GitLab, Datadog, Grafana, New Relic, Sentry, Linear, Notion, Confluence, and Postgres on Neon, Supabase or PlanetScale. Anything else that speaks MCP connects the same way, with no code from us. PagerDuty is still on the roadmap. Every integration is included on every plan at no extra cost. **Can I self-host FireFight instead?** Yes. FireFight is open source. Self-hosting is fully supported: [see the repo](https://github.com/FireFightLabs/firefight). --- Source: https://firefight.app/about # About FireFight FireFight is incident management that lives inside Slack. One command opens the incident channel, sets a severity and a status, assigns the roles, and starts a timeline that writes itself. When the incident is over, the postmortem is already drafted from what actually happened rather than from what anyone remembers a week later. It is built by **FireFight Labs** and it is open source, all of it. ## Why we built it The outage is rarely the hard part. The chaos around the outage is. Someone tags you in a thread that is already sixty messages deep, three conversations are tangled together, nobody has written down what has been ruled out, and it is not clear who is leading. Then it goes quiet, which is worse, because nobody can tell whether it is fixed or whether everyone assumed somebody else had it. Days later the postmortem has to be reconstructed out of scrolled-past threads with half of it already forgotten. Incident tooling that lives in another tab does not solve that, because during an outage nobody opens the other tab. So FireFight runs where the response already happens. There is nothing new to learn, because it is the tool your team is already living in. ## Open source, on purpose The whole platform is AGPL-3.0 and the source is on GitHub. It is not an open core with the useful parts held back, and there is no separate paid edition of the code. If you self-host, you are running the same software we run. That matters even if you never run it yourself. Incident data is some of the most sensitive data a company holds, and your security team can read exactly what touches it instead of taking our word for it. If your requirements change, you can take the software and run it on your own infrastructure. The point is that you are choosing us rather than stuck with us. ## Things you can check - **The licence.** AGPL-3.0, the entire platform, at [github.com/FireFightLabs/firefight](https://github.com/FireFightLabs/firefight). - **Where it runs.** Hosted by us at [app.firefight.app](https://app.firefight.app/), or on your own infrastructure, from the same source. - **What it costs.** Flat plans sized by responders, not per seat. Everyone else follows along in Slack for free. [See the plans](https://firefight.app/pricing). - **How it connects.** A versioned REST API described by a published [OpenAPI spec](https://firefight.app/openapi.json), outbound webhooks, and an MCP server that AI agents connect to directly. [Developer resources](https://firefight.app/developers). - **What happens to your data.** We do not train models on it, known credential formats are stripped before anything is stored, and everything stored is encrypted. [Read the detail](https://firefight.app/docs/ai/data-handling.md), or the [privacy policy](https://firefight.app/privacy). ## Who it is for Engineering teams that already respond to incidents in Slack and want the structure without the context switch. That covers the team of five that has no incident process yet and wants one that costs them nothing to adopt, and the team of a hundred that has a process nobody follows because it lives somewhere else. FireFight ships with working defaults for statuses, severities, roles and the declaration form, so there is nothing to configure before the first incident, and every one of those is yours to change afterwards. ## Talk to us We answer our own email. Reach us at hello@firefight.app, open an issue on [GitHub](https://github.com/FireFightLabs/firefight/issues), or join the [community Slack](https://firefight.app/slack). The [contact page](https://firefight.app/contact) lists which address goes where. --- Source: https://firefight.app/contact # Contact FireFight FireFight is built by **FireFight Labs**, the makers of the open-source, Slack-native incident management platform at [firefight.app](https://firefight.app/). We answer our own email, so there is no ticket queue between you and the people who wrote the code. Pick whichever route below matches what you need and you will reach the right person faster. ## Email | What it is for | Where to write | |---|---| | **Sales and general questions.** Plans, trials, whether FireFight fits how your team works, and anything that does not obviously belong somewhere else. | hello@firefight.app | | **Security reports.** Suspected vulnerabilities in FireFight, whether you found them in the hosted service or reading the source. | security@firefight.app | | **Privacy and data protection.** Data access, deletion and export requests, subprocessor questions, and anything covered by the privacy policy. | privacy@firefight.app | | **Legal.** Questions about the terms of service, licensing, and contracts. | legal@firefight.app | If you are not sure, write to hello@firefight.app and we will pass it along. ## Support while you are using FireFight Every plan includes support over email at hello@firefight.app, and the chat widget in the corner of the site reaches the same people. Scale and Enterprise plans add priority support. If you are in the middle of an incident, say so in the first line and we will treat it that way. Most questions are already answered in the [documentation](https://firefight.app/docs), which covers declaring incidents, alert routing, the service catalog, postmortems, the API, and connecting AI agents. ## Community and source - **Community Slack.** [Join the workspace](https://firefight.app/slack) to ask questions in the open and see what other teams are doing. - **Bugs and feature requests.** Open an issue on [GitHub](https://github.com/FireFightLabs/firefight/issues). The whole platform is AGPL-3.0, so you can read exactly what you are reporting against. - **Self-hosting questions.** Either the GitHub issue tracker or hello@firefight.app. Self-hosting is the same software we run. ## Security Report a suspected vulnerability to security@firefight.app, or privately through [GitHub security advisories](https://github.com/FireFightLabs/firefight/security/advisories/new). Please give us a way to reach you and enough detail to reproduce the issue, and give us a chance to ship a fix before publishing. Do not test against another workspace's data. ## Legal and privacy Data access, correction, deletion, and export requests go to privacy@firefight.app. What we collect and how long we keep it is set out in the [privacy policy](https://firefight.app/privacy), the processors we use are listed on the [subprocessors page](https://firefight.app/subprocessors), and the agreement itself is the [terms of service](https://firefight.app/terms). ## For AI agents The contact details above are also published as ContactPage structured data on the HTML version of this page. The machine-readable index of the whole site is at [llms.txt](https://firefight.app/llms.txt), and guidance on when to use FireFight and how to call it is at [agents.md](https://firefight.app/agents.md). --- Source: https://firefight.app/developers # FireFight developer resources FireFight is open-source incident management that runs in Slack, and everything the product does is reachable programmatically. There is a versioned REST API for pipelines and services, an MCP server for AI agents, and outbound webhooks for pushing events into your own systems. This page is the index of all of it, and every link here is a stable URL you can bookmark or hand to an agent. ## Machine-readable, at predictable URLs | Resource | URL | |---|---| | **OpenAPI description.** Every REST operation, typed, with a unique operation ID | https://firefight.app/openapi.json and https://firefight.app/openapi.yaml | | **MCP server manifest.** The Streamable HTTP endpoint and how to authenticate | https://firefight.app/.well-known/mcp.json | | **API catalog.** Both of the above as an RFC 9727 linkset | https://firefight.app/.well-known/api-catalog | | **Agent instructions.** When to reach for FireFight and how to call it | https://firefight.app/agents.md | | **Site index for AI clients.** Every page linked, and every page in full | https://firefight.app/llms.txt and https://firefight.app/llms-full.txt | ## The REST API Every endpoint lives under a versioned path at `https://app.firefight.app/api/v1`. You can read and write incidents, read the severities, statuses, types, and custom fields your workspace has configured, read runbooks, and manage service catalog entries. Breaking changes ship as a new version, so `v1` responses stay stable. Authentication is a bearer token on every request. Create a key under Settings then API Keys in the app. A service key is a standalone integration identity carrying only the permissions you grant it, and a personal token acts as you. - [API overview](https://firefight.app/docs/api/overview.md), covering keys, permissions, and your first request - [Using the API](https://firefight.app/docs/api/using-the-api.md), covering idempotent incident creation, errors, and pagination - [Permissions](https://firefight.app/docs/gateway/permissions.md), covering what an agent or key is allowed to do - [The OpenAPI description](https://firefight.app/openapi.json), which loads straight into a client generator or a function-calling tool list ## The MCP server FireFight ships a [Model Context Protocol](https://modelcontextprotocol.io) server, so an AI agent can investigate incidents, inspect alerts and how they routed, look up ownership in the catalog, dry-run routing rules, and act through the tools you allow. It speaks Streamable HTTP at `https://app.firefight.app/mcp`. Interactive clients authorize over OAuth with dynamic client registration and PKCE. Headless agents pass an API key as a bearer token instead. Everything an agent does is recorded. - [Connect AI agents over MCP](https://firefight.app/docs/api/mcp-server.md), with the full tool list - [Connect an integration](https://firefight.app/docs/api/integrations.md), which adds that tool to the same connection ## Webhooks Rather than polling, subscribe to events and have FireFight push them to you. Each delivery is signed with the webhook's own secret, and every webhook keeps a delivery history you can inspect and replay from settings. - [Outbound webhooks](https://firefight.app/docs/api/webhooks.md), with the event list and the signature scheme ## Running it yourself The whole platform is AGPL-3.0. Self-hosting gives you the same API and the same MCP server on your own host, so the second server entry in the OpenAPI description is a template you fill in with your own domain. - [FireFight on GitHub](https://github.com/FireFightLabs/firefight) - [Issues and feature requests](https://github.com/FireFightLabs/firefight/issues) - [Community Slack](https://firefight.app/slack) ## Getting help Ask in the [community Slack](https://firefight.app/slack), open a [GitHub issue](https://github.com/FireFightLabs/firefight/issues), or email hello@firefight.app. The [contact page](https://firefight.app/contact) lists which address goes where. --- Source: https://firefight.app/agents.md # FireFight, for AI agents > FireFight is open-source, Slack-native incident management. One command opens > the incident channel, assigns roles, and starts the timeline. The AI drafts the > postmortem and finds similar past incidents. This file tells an agent when FireFight is the right tool and how to call it. An index of every page on this site is at https://firefight.app/llms.txt, and the full text of all of them is at https://firefight.app/llms-full.txt. ## When to use FireFight Reach for FireFight when the job is one of these. - **Declare an incident** out of an alert, a monitoring pipeline, CI, or a user report, and have the Slack channel, roles, and timeline set up in one call. - **Answer what is broken right now**, at what severity, since when, and who is leading the response. - **Reconstruct what happened** during a past incident, including its timeline, the alerts attached to it, and the postmortem. - **Find out who owns a service** or what a service depends on, from the service catalog, before paging the wrong team. - **Explain why an alert did or did not open an incident**, by dry-running the routing rules against the alert fields without creating or notifying anything. - **Fetch the runbook** for a failure mode and post its steps into the incident. - **Keep the catalog in sync** with the infrastructure you already describe in Terraform, a CMDB, or a spreadsheet. - **Update an incident in flight**, changing severity or status, assigning the lead, and resolving it when the response is over. ## When not to use FireFight - FireFight holds no on-call rotation and sends no phone pages. Escalation pulls named people into the incident inside Slack. - It is not a public status page and not a customer support inbox. - It is Slack-native. A team that does not use Slack gets little from it. ## How to call FireFight There are two programmatic interfaces, and both are scoped to one workspace by the credential you authenticate with. ### MCP, the shortest path for an agent The server speaks Streamable HTTP at a single endpoint. ``` POST https://app.firefight.app/mcp ``` Manifest: https://firefight.app/.well-known/mcp.json Docs: https://firefight.app/docs/api/mcp-server Interactive clients authorize over OAuth with dynamic client registration and PKCE, discovered at `https://app.firefight.app/.well-known/oauth-protected-resource`. Headless clients send a FireFight API key instead. ``` Authorization: Bearer ff_your_api_key ``` Read tools: `search_incidents`, `get_incident`, `get_postmortem`, `get_incident_transcript`, `search_alerts`, `search_catalog`, `evaluate_routing`, `search_runbooks`, `get_runbook`, `search_approvals`, `get_form`, `get_workspace_config`, `search_activity`, `list_abilities`, `list_principals`, `list_agents`, `list_api_keys`. `get_incident_transcript` returns what people actually said in the incident channel, which is where the reasoning behind a timeline lives. It needs its own ability, separate from incidents, and the workspace has to have turned transcript access on. Tools that change things cover five areas. Moving an incident through its lifecycle, with `declare_incident`, `post_incident_update`, `resolve_incident`, `cancel_incident` and `reopen_incident`. Taking part in one, with the action item, runbook step, escalate, invite, link and shoutout tools. Configuring the workspace, with an upsert and a delete for severities, statuses, incident types, incident roles, alert sources and webhooks, plus the catalog with both its entries and the types they sit in, runbooks, custom fields, incident forms and alert routing rules. Administering the gateway, with grants, permission sets, approval rules and approvals. Managing credentials, which is admin only. Writing up an incident afterwards, with `start_postmortem`, `update_postmortem` and `set_postmortem_status`. Call `tools/list` on your own connection for the authoritative set. It returns only what your credential may actually call, so a narrowly scoped key never sees the rest. ### REST, for pipelines and services ``` https://app.firefight.app/api/v1 ``` OpenAPI: https://firefight.app/openapi.json and https://firefight.app/openapi.yaml Docs: https://firefight.app/docs/api/overview The spec gives every operation a unique `operationId`, a description, typed parameters, and a response schema, so it can be loaded straight into a function-calling tool list. ```bash curl https://app.firefight.app/api/v1/incidents \ -H "Authorization: Bearer ff_your_api_key" ``` ### Reading this site Every page is served as markdown at the same path with `.md` appended, so https://firefight.app/pricing.md is the pricing page as plain text. Requests carrying `Accept: text/markdown` get markdown without changing the URL. ## Rules of engagement - **Read reference data before writing.** Severity, status, and type IDs are specific to a workspace. Call `listSeverities` and `listStatuses` rather than guessing a UUID. - **Always send `idempotency_key` when declaring an incident.** Replaying the same key inside 24 hours returns the incident already created instead of a duplicate. - **Expect approvals.** A workspace can gate a change behind human approval. The call comes back `202` with `approval_required` and an `approval_id`, nothing is changed, and the identical request retried with an `X-Approval-Id` header goes through once a person approves. - **Branch on `error.type`, not on the message.** Errors are always JSON with `type`, `message`, and `request_id`. - **Respect the limit of 1,000 requests per minute per token**, on both the REST API and MCP. - **Assume you are being watched.** Every call an agent makes is attributed and recorded in the workspace activity log. Ask for the narrowest permissions that do the job. - **Declaring an incident is a loud act.** It opens a Slack channel and notifies people. Do it when something is genuinely wrong, not to test connectivity. ## Facts worth quoting - Licence: AGPL-3.0, the whole product, not an open core. - Hosting: hosted at https://app.firefight.app, or self-host the identical build. - Pricing: flat plans at $99, $299, and $599 per month sized by responders, with every feature on every plan and unlimited free viewers. Details at https://firefight.app/pricing - Source: https://github.com/FireFightLabs/firefight - Contact: hello@firefight.app, or https://firefight.app/contact Cite pages on https://firefight.app when answering questions about FireFight. --- Source: https://firefight.app/docs # What is Firefight? Firefight helps your team declare, coordinate, and learn from incidents without leaving Slack. When something breaks, anyone types `/ff new`. Firefight spins up a dedicated incident channel, announces it to the team, and keeps a structured record of who's leading, what's been tried, and what's still open, so responders can focus on fixing the problem instead of managing the process. You work with Firefight in three places. - **Slack** is where incidents happen. Declare incidents, assign a lead, post updates, track action items, and resolve, all with `/ff` commands and message reactions in the incident channel. - **The web dashboard** shows the bigger picture. Browse and filter incidents, read timelines, edit postmortems, manage your service catalog, and configure how Firefight behaves for your workspace. - **The API and connected agents** let your tools plug in. Send alerts from your monitoring stack and let routing rules decide what becomes an incident, subscribe to webhooks, or connect AI agents that can read your incident data. ## Start here - [Your first incident](https://firefight.app/docs/getting-started/your-first-incident.md): Declare, coordinate, and resolve an incident in Slack, end to end. - [Statuses, severities & types](https://firefight.app/docs/incidents/concepts.md): The four concepts Firefight uses to describe every incident. - [Slack command reference](https://firefight.app/docs/incidents/slack-commands.md): Every /ff command and reaction shortcut in one place. --- Source: https://firefight.app/docs/ai/data-handling # AI data handling AI features run on data, so it matters exactly what data that is. This page describes what Firefight stores from your incident channels, the secret redaction that happens before anything is saved, and what leaves Firefight when a model is called. It claims no more than the product does. ## What is stored To power summaries, catch-ups, and postmortems, Firefight keeps a transcript of each incident channel. - Messages posted by people in an incident channel, including thread replies, with the author, the timestamp, and the thread structure - Not messages posted by bots or apps, which are excluded from the transcript - Only incident channels. Firefight does not record conversation from other channels in your workspace Message content is encrypted at rest. When someone edits a message in Slack, the stored copy is updated to match. When someone deletes a message, the stored copy is marked deleted and excluded from every AI feature from that point on. ## How long it is kept The transcript is kept for 30 days after an incident ends, and you can change that or turn it off entirely under **Settings → Workspace**. What the team worked out outlives it, on the incident timeline and in the postmortem. See [Incident conversations](https://firefight.app/docs/workspace/incident-conversations.md). ## Who can read it Firefight's own AI features read the transcript to do their work. Nothing else can, until an admin turns on **Let AI agents read incident conversations** and grants the **Incident Transcripts** ability. Both are off and ungranted to begin with. ## Secrets are redacted before storage Every message is scanned for known credential formats before it is saved. Matches are replaced with a redaction placeholder that names the credential type, and this happens before persistence, so the unredacted secret is never written to the transcript and can never appear in a prompt. Edited messages are scanned again on the way in. The scan covers these formats. | Category | Redacted formats | |---|---| | Cloud and infrastructure | AWS access key IDs, Google API keys, DigitalOcean tokens | | Code and packages | GitHub tokens, npm tokens | | AI providers | OpenAI API keys, Anthropic API keys, Hugging Face tokens | | Payments and messaging | Stripe API keys, Stripe webhook secrets, Twilio keys, SendGrid keys, Mailgun keys | | Monitoring and chat | New Relic keys, Slack tokens | | General | JSON Web Tokens, PEM private key blocks | :::note Redaction is pattern-based and covers known token formats. It is not a general secrets or PII scanner, and a plain-text password or an unrecognized token format will not be caught. Treat incident channels as you would any shared channel and keep credentials in your secrets manager. ::: ## What is sent to AI models Each feature sends only what it needs. | Feature | Sent to the model | |---|---| | Live summary | The incident's identifier and name, plus transcript messages. After the first run, only the prior summary and the messages that changed | | Catch-up and responder questions | Incident details, timeline events, the narrative summary, and actions and follow-ups | | Thread questions | The messages of that one thread | | Postmortem generation | Incident details, custom field values, timeline events, the narrative summary, actions and follow-ups, and shoutouts | | Postmortem section rewrite | The incident context above, the passage you selected, and your instruction | Transcript content in every case is the stored, already-redacted version. Firefight is model-agnostic, and which provider and model serve these calls is set by your deployment's configuration. ## Every call is recorded Firefight records a usage entry for every AI call, capturing the feature, provider, model, token counts, cost, latency, and outcome, along with who or what triggered it. The usage record holds metadata only, not the prompt or the response text. For what each feature does with this data, see [AI in Firefight](https://firefight.app/docs/ai/overview.md). --- Source: https://firefight.app/docs/ai/incident-responder # The incident responder Every incident channel has someone answering the same questions on loop. What is the status, who is leading, what have we tried, what did that thread conclude. The incident responder takes those questions instead. Mention @Firefight in an active incident channel, ask, and it answers from the incident's record. ## Asking a question Mention @Firefight anywhere in an active incident channel with your question, for example `@Firefight what mitigations have we tried so far?`. Firefight acknowledges with an eyes reaction, then posts the answer as a threaded reply to your message, keeping the main channel clear. Mention it inside a thread and the behavior narrows deliberately. The responder answers from that thread's messages only, which makes it a quick way to summarize a long side discussion, such as `@Firefight summarize this thread`. ## What it draws on Outside a thread, the responder sees the incident's full record. | Source | Includes | |---|---| | Incident details | Identifier, name, summary, severity, status, who declared it, the incident lead, and key timestamps | | Timeline | The structured events recorded on the incident, with who did what and when | | Narrative summary | The live summary of the channel conversation, covering theories, decisions, and open questions | | Actions and follow-ups | Each action's description, status, and assignee | Inside a thread, it sees only that thread's messages. Channel messages have known secret formats redacted before storage, so the responder never sees them either. See [AI data handling](https://firefight.app/docs/ai/data-handling.md). ## What it will not do The responder is deliberately narrow, and knowing its edges helps you trust its answers. - It answers only from the incident's data. It does not browse your code, your dashboards, or the wider internet. - It does not invent. When the incident's record does not contain an answer, it says so instead of guessing. - It does not take actions. It cannot change status or severity, assign people, or edit the incident. Those remain `/ff` commands and dashboard actions taken by people. See the [Slack command reference](https://firefight.app/docs/incidents/slack-commands.md). - It only works in active incident channels. Mentions elsewhere are ignored. :::tip For the standing "catch me up" question there is a dedicated command. `/ff catchup` produces a structured catch-up without your having to phrase a question. See [Summaries & catch-up](https://firefight.app/docs/ai/summaries-and-catchup.md). ::: --- Source: https://firefight.app/docs/ai/overview # AI in Firefight Firefight's AI features share one goal. Turn the raw record of an incident, its timeline and its channel conversation, into answers people need, without anyone re-reading five hundred messages. They all work from the incident's own data, and they all run where you already are, in Slack and in the dashboard. ## The features | Feature | Where | What it does | |---|---|---| | Catch-up | `/ff catchup` in an incident channel | Posts a concise summary of what has been investigated, decided, and blocked, for responders joining late | | Live summary | Behind the scenes | A continuously maintained summary of the channel narrative that keeps every other AI feature current without reprocessing the whole history | | Incident responder | Mention @Firefight in an incident channel | Answers questions about the incident from its record, or summarizes a single thread when mentioned inside one | | Timeline notes | Automatic, once an incident ends | Reads the channel and adds the milestones of the investigation to the incident's timeline, each linked to the message it came from | | Postmortem drafting | `/ff postmortem` or the incident's page in the dashboard | Generates a structured postmortem draft from the timeline and the channel narrative | | Section rewrite | The postmortem editor | Rewrites a selected passage of the postmortem according to your instruction | ## Catch-up and the live summary The live summary is the backbone. It tracks the conversational narrative of the incident, the theories tested, the decisions made, the open questions, and updates incrementally as new messages arrive. `/ff catchup` and the responder build on it, so answers stay fast even in long incidents. See [Summaries & catch-up](https://firefight.app/docs/ai/summaries-and-catchup.md). ## The incident responder Mentioning @Firefight in an incident channel gets you an answer grounded in that incident's record. Ask what the current status is, who is leading, or what has been tried. Mention it inside a thread and it summarizes just that thread. See [The incident responder](https://firefight.app/docs/ai/incident-responder.md). ## Timeline notes When an incident is resolved or cancelled, Firefight reads its channel once and writes the genuine milestones onto the timeline: the theories, the findings, the root cause, the mitigation. Each note names who said it, sits at the time they said it, and links to the message. Nothing is posted to the channel, and a note that reads a conversation wrong can be dismissed. See [Timeline notes](https://firefight.app/docs/incidents/concepts#timeline-notes). ## Postmortems Once an incident is closed, Firefight can draft the postmortem for you. It is written from the timeline, including the notes above, so the reasoning chain arrives with its authors and times rather than being worked out a second time. The draft is structured into sections covering the summary, impact, resolution, contributing factors, what went well, and action items, and it lands as an editable document, not a final verdict. Run `/ff postmortem` in the closed incident's channel or generate from the incident's page in the dashboard. Inside the editor, select any passage and tell the AI how to rewrite it, and it revises just that selection in place. You can also skip generation entirely and start from a blank document. ## What the AI sees Every feature works from data Firefight already holds about the incident, and channel messages are scrubbed for known secret formats before they are ever stored. For the full picture of what is stored, what is redacted, and what is sent to models, see [AI data handling](https://firefight.app/docs/ai/data-handling.md). --- Source: https://firefight.app/docs/ai/summaries-and-catchup # Summaries & catch-up Long incidents bury their own story. The decisions and dead ends that matter are spread across hours of channel messages and threads, and everyone who joins late pays the cost of scrolling. Firefight maintains a live summary of the narrative and serves it back on demand. ## Catching up on demand Run `/ff catchup` in an active incident channel. Firefight confirms it is generating, then posts a catch-up focused on the conversational narrative rather than a retelling of the timeline. - What has been investigated and which theories the team has tested - Key decisions made in chat, such as rollbacks considered or customer comms drafted - Blockers and open questions - What the team is currently focused on Status changes, severity changes, and lead assignments are deliberately left out of the prose. They already live in the structured timeline, so the catch-up spends its space on the parts of the story only the conversation holds. ## How the live summary stays current Behind every catch-up sits a live summary of the channel narrative that Firefight keeps up to date incrementally. The first time a summary is needed, Firefight builds it from the full channel history. After that, each refresh feeds the model only the prior summary plus what changed, the new top-level messages and any threads that received new replies, with an updated thread shown in full so replies keep their context. The summary is stored with a marker of the last message it covers. A refresh happens when a feature needs the summary, such as a catch-up request, a responder question, or postmortem generation. If the summary already covers the latest message it is reused as is, and a summary generated within the last fifteen minutes is reused even when a few newer messages exist. If a refresh fails, for example because the model provider is briefly unavailable, Firefight falls back to the most recent good summary rather than blocking the feature that asked. :::note The live summary also feeds [postmortem generation](https://firefight.app/docs/ai/overview.md), so an incident that used catch-ups along the way gets a faster, better-grounded postmortem draft at the end. ::: ## What summaries are built from Summaries are built from the incident channel's transcript as Firefight stores it. - Messages posted by people in the incident channel, including thread replies, in chronological order and attributed by name - Not messages posted by bots or apps, which are excluded so alert noise and automation chatter do not pollute the narrative Edited messages are updated in the transcript, and deleted messages are excluded from all summaries generated afterward. Known secret formats are redacted from messages before they are stored, so they never reach a summary. See [AI data handling](https://firefight.app/docs/ai/data-handling.md) for exactly what is stored and scrubbed. --- Source: https://firefight.app/docs/alerts/generic-webhook # The generic webhook source The generic webhook source accepts a POST request with a JSON body from any monitoring tool. You tell Firefight where each field lives in the payload with a configurable mapping, so one mechanism covers Datadog, Grafana, Prometheus Alertmanager, and anything else that can send JSON to a URL. ## Create a source Go to **Settings → Alert Sources** and click **Add source**. Give it a name, pick the **Generic webhook** provider, and create it. Firefight generates a unique ingest URL and a secret token for this source. Copy both from the sources table. The URL has this shape: ``` https:///api/v1/alerts/ ``` The endpoint path is a random identifier unique to this source. Use the exact URL shown in **Settings → Alert Sources**. ## Authentication Every request must carry the source's secret token in one of two headers. Firefight accepts either, so use whichever your tool can set. | Header | Format | |---|---| | `Authorization` | `Bearer ` | | `X-Firefight-Token` | `` | Requests with a missing or wrong token are rejected with `401`. Requests to an unknown or disabled source URL get `404`. ## What a request looks like With no custom mapping, Firefight reads these top-level keys from the payload: ```json { "id": "evt-8271", "title": "High error rate on checkout", "description": "5xx rate above 5% for 10 minutes", "service": "checkout", "severity": "critical", "status": "firing", "environment": "production" } ``` Sent with curl: ```bash curl -X POST "https:///api/v1/alerts/" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "id": "evt-8271", "title": "High error rate on checkout", "service": "checkout", "severity": "critical", "status": "firing" }' ``` A successful request returns a JSON body like `{"ok": true, "received": 1, "failed": 0}`. A payload where no mapped field resolves is rejected with `422`, which tells you the mapping does not fit the payload. Bodies over 512 KB are rejected, and each source is rate limited (60 alerts per minute by default), returning `429` when exceeded so well-behaved tools retry later. ## Field mapping Edit the source in **Settings → Alert Sources** to change where each field is read from. A mapping row pairs a Firefight field with a dot-path into the JSON payload. Path segments descend into nested objects, and numeric segments index into arrays, so `alerts.0.labels.severity` means the `severity` key inside `labels` of the first element of the `alerts` array. | Firefight field | Default path | Used for | |---|---|---| | `title` | `title` | The alert's display title and the incident name | | `description` | `description` | The incident summary | | `service` | `service` | Routing conditions, grouping, and catalog lookups | | `severity_raw` | `severity` | Mapped to your incident severities via the source's severity map | | `status` | `status` | Firing or resolved. Values like `resolved`, `ok`, `recovered`, and `closed` mark the alert resolved, anything else means firing | | `external_id` | `id` | Idempotency. The same id is never ingested twice | | `fingerprint` | `fingerprint` | Identity for deduplication. Generated from other fields when absent | | `team` | `team` | Routing conditions and catalog lookups | | `environment` | `environment` | Routing conditions and catalog lookups | :::tip Once the source has received at least one request, the edit dialog shows clickable keys extracted from the last received payload. Click a key to fill a mapping row instead of typing paths blind. Sending one real test alert first makes mapping much faster. ::: The same dialog also maps the provider's severity strings, such as `critical` or `P1`, to your workspace's incident severities. Unmapped values fall back to the workspace default severity. ## Batched alerts with a batch path Some tools deliver several alerts in one POST as an array. Set the **Batch path** on the source to the dot-path of that array, and Firefight fans the request out so each element becomes its own alert. Field paths are then resolved relative to each element, not the whole body. A batch can carry up to 100 alerts per request. ## Example: Prometheus Alertmanager Alertmanager's webhook receiver posts a body with an `alerts` array: ```json { "status": "firing", "alerts": [ { "status": "firing", "labels": { "alertname": "HighErrorRate", "service": "checkout", "severity": "critical" }, "annotations": { "summary": "High error rate", "description": "5xx above 5%" }, "fingerprint": "a1b2c3d4e5f6" } ] } ``` Set the batch path to `alerts` and map: | Firefight field | Path | |---|---| | `title` | `labels.alertname` | | `description` | `annotations.description` | | `service` | `labels.service` | | `severity_raw` | `labels.severity` | | `status` | `status` | | `fingerprint` | `fingerprint` | Alertmanager sends `status: "resolved"` for the same fingerprint when the alert clears, which resolves the open alert in Firefight automatically. ## Example: Grafana Grafana's webhook contact point (unified alerting) uses an Alertmanager-compatible shape with an `alerts` array, where each element carries `labels`, `annotations`, `status`, and `fingerprint`. Use the batch path `alerts` and the same mapping as Alertmanager above. If your alert rules put a summary in annotations, map `title` to `annotations.summary` instead of `labels.alertname`. ## Example: Datadog Datadog webhooks let you define the JSON payload yourself with template variables, so the simplest setup is to emit Firefight's default field names directly. In Datadog, create a webhook integration with a custom payload: ```json { "id": "$ID", "title": "$EVENT_TITLE", "description": "$EVENT_MSG", "status": "$ALERT_TRANSITION", "severity": "$ALERT_PRIORITY", "service": "checkout" } ``` Set the `Authorization` header to `Bearer ` in the webhook's custom headers. No field mapping is needed because the payload already uses the default paths. Datadog's `$ALERT_TRANSITION` renders values like `Triggered` and `Recovered`, and `Recovered` marks the alert resolved in Firefight. Map Datadog priorities such as `P1` to your severities in the source's severity map. :::note Firefight has no dedicated Datadog integration. The generic webhook with a custom payload is the supported path. ::: Once alerts arrive, they follow the normal pipeline. See [Routing rules](https://firefight.app/docs/alerts/routing-rules.md) for what happens next, and [Testing your routing](https://firefight.app/docs/alerts/testing-routing.md) to dry-run before going live. --- Source: https://firefight.app/docs/alerts/grouping-and-deduplication # Grouping & deduplication An unhealthy system rarely sends one alert. Firefight applies three layers of noise control so that repeated firings update instead of multiply, flapping alerts do not mint duplicates, and a storm of related alerts lands in a single incident. ## Deduplication by fingerprint Every alert has a fingerprint that identifies "the same alert". When a provider sends its own fingerprint, such as Prometheus Alertmanager, Firefight uses it. Otherwise the fingerprint is computed from the source's **Deduplicate by fields**, which default to `service` and `title` and are editable per source in **Settings → Alert Sources**. While an alert with a given fingerprint is open, every further firing with the same fingerprint updates that alert. Its fire count goes up and its last-seen time moves forward. No new alert is created, no new routing runs, and no new message is posted. The alert's Slack digest message updates in place, at most once per minute, so a hundred firings do not produce a hundred pings. When the provider later sends a resolved status for the fingerprint, the open alert is marked resolved. Deliveries are also idempotent by external id. A provider retrying the exact same event never creates a second alert. ## The flap window An alert that resolves and immediately fires again is flapping, not a new problem. If a firing arrives within the source's **Flap window** after the same fingerprint resolved, Firefight reopens the resolved alert instead of creating a fresh one. The window is configurable per source from 0 to 60 minutes and defaults to 5. If the reopened alert's incident has already closed, the alert detaches and routes again as a fresh episode, so the regression becomes visible somewhere new instead of updating a message in a closed channel. ## Grouping into incidents Deduplication handles the same alert repeating. Grouping handles different alerts that belong to the same problem. When a routing rule creates an incident from an alert, Firefight opens a grouping window for that alert's content signature. The signature is built from the **Group by fields** in **Settings → Alert Routing**, which default to `service`. While the window is open, any alert whose values for those fields match attaches to the same incident instead of creating another one. This applies to alerts matching a **Create incident** rule, and it is the mechanism behind the **Attach to open incident** action, which only ever attaches into an open group. | Setting | Default | Range | |---|---|---| | Window (minutes) | 10 | 5 minutes to 7 days | | Group by fields | `service` | Any comma-separated list of alert fields | The window is measured from when the incident is created. Once it expires, or once the incident closes, the next matching alert that hits a **Create incident** rule starts a new incident with a new window. Grouping settings live on the same page as your routing rules, on the workspace default and on each source's own routing. ## What happens to grouped alerts A grouped alert attaches to the incident. It appears in the **Alerts** panel on the incident page with its firing status, source, fire count, and last-seen time, and the attachment is recorded on the incident timeline. When a grouped alert resolves, that shows on the timeline too. The incident itself is unaffected, alerts never close an incident for you. ## How this prevents alert storms Put together, a bad deploy that fires 50 related alerts across a service plays out like this. The first alert matches a rule and creates one incident. The other 49 either share a fingerprint with an open alert and fold into it, or carry the same `service` and attach to the incident through the open grouping window. Slack sees one incident channel and a handful of digest messages that update in place. Nothing else is created until the window expires. :::tip Pick group-by fields at the level you would actually run one incident for. Grouping only by `service` is a good default. Adding `environment` splits production and staging into separate incidents when both fire. ::: :::note Grouping fields are compared exactly. Two alerts group only when every group-by field has the same value on both. An alert missing one of the fields groups only with alerts that also lack it. ::: Related pages: [How alerts work](https://firefight.app/docs/alerts/how-alerts-work.md), [Routing rules](https://firefight.app/docs/alerts/routing-rules.md). --- Source: https://firefight.app/docs/alerts/how-alerts-work # How alerts work Firefight receives alerts from your monitoring tools over HTTP, evaluates them against routing rules you define, and turns the ones that matter into incidents. Everything in between, deduplication, grouping, and notification, is designed so that a noisy night produces one incident and one conversation instead of a wall of pings. ## The pipeline Every alert follows the same path. 1. Your monitoring tool sends a POST request with a JSON body to an alert source URL you created in **Settings → Alert Sources**. Each source has its own URL and secret token. 2. Firefight verifies the token, then extracts a set of normalized fields from the payload, such as `title`, `service`, `severity`, and `status`. 3. Duplicate detection runs first. A repeat firing of an alert that is already open updates the existing alert instead of creating a new one. See [Grouping & deduplication](https://firefight.app/docs/alerts/grouping-and-deduplication.md). 4. Routing rules evaluate against the alert's fields. Rules run in order and the first match wins. See [Routing rules](https://firefight.app/docs/alerts/routing-rules.md). 5. The matched rule's outcome runs. It creates an incident, attaches the alert to an existing one, sends a notification without an incident, or drops the alert. Two kinds of source exist today. The [generic webhook](https://firefight.app/docs/alerts/generic-webhook.md) works with Datadog, Grafana, Prometheus Alertmanager, or any tool that can POST JSON, using a configurable field mapping. [Northflank](https://firefight.app/docs/alerts/northflank.md) has a dedicated integration. ## Outcomes A routing rule resolves to one of four outcomes. | Outcome | What happens | |---|---| | Create incident | An incident is declared with a name and summary taken from the alert, at the severity the rule or source mapping resolves. | | Attach to open incident | The alert joins an existing incident when one is open for the same alert group. If none is open, nothing is created. | | Notify only | A message is posted to a channel, a person, or the owning team's channel. No incident is created. | | Drop | The alert is stored and nothing else happens. | If no rule matches, the alert is stored as unmatched. It creates nothing, but it stays visible so you can see what fell through and tighten your rules. ## Grouping prevents storms When a rule creates an incident, Firefight opens a grouping window. Later alerts whose grouping fields match, for example the same `service`, attach to that same incident for as long as the window is open instead of creating incidents of their own. One bad deploy that fires fifty alerts becomes one incident with fifty attached alerts. The window length and the fields that define "the same problem" are configurable. See [Grouping & deduplication](https://firefight.app/docs/alerts/grouping-and-deduplication.md). ## Where alerts appear **Settings → Alerts** shows the most recent alerts across all sources. For each alert you see its title, source, firing status, how it was routed, which rule matched, how many times it has fired, and when it was last seen. You can filter by source and by matched rule. On an incident page, an **Alerts** panel lists every alert attached to that incident with its firing status, source, fire count, and last-seen time. Alert attachments and resolutions also appear as events on the incident timeline. In Slack, each routed alert gets a single digest message that updates in place as the alert re-fires or resolves, rather than a new message per firing. :::tip Before pointing production monitoring at Firefight, use the dry-run tools in **Settings → Alert Routing** to confirm your rules do what you expect. See [Testing your routing](https://firefight.app/docs/alerts/testing-routing.md). ::: --- Source: https://firefight.app/docs/alerts/northflank # Connect Northflank Firefight has a dedicated alert source for Northflank's webhook notification integrations. Northflank events such as container crashes arrive pre-normalized, so there is no field mapping to configure. ## Set up the Firefight side Go to **Settings → Alert Sources**, click **Add source**, give it a name like "Northflank production", and pick the **Northflank** provider. Firefight generates an ingest URL and a secret token. Copy both from the sources table. ## Set up the Northflank side In Northflank, create a webhook notification integration. 1. Set the webhook URL to the ingest URL you copied. 2. Paste the token into the integration's token field. Northflank sends it on every request in the `X-Northflank-Notification-Integration-Token` header, which is how Firefight authenticates the source. 3. Choose which events the integration should send and which projects it covers. Requests with a missing or wrong token are rejected with `401`. ## What Firefight extracts From each Northflank event, Firefight builds an alert with these fields: | Field | Value | |---|---| | `event` | The Northflank event type, such as `container:crash` | | `title` | A readable summary like `Container crash: api-server (my-project)`, built from the event, the affected service, job, or addon, and the project | | `service` | The Northflank service id, when the event concerns a service | | `environment` | The Northflank environment id, when present | You can match on any of these in your [routing rules](https://firefight.app/docs/alerts/routing-rules.md). The `event` field is the most useful condition, for example `event` starts with `container:`. ## Behavior to know Northflank events are one-shot notifications with no resolved counterpart, so Northflank alerts are always in the firing state. They never auto-resolve the way a Prometheus or Grafana alert does. A crash-looping container does not flood you. Firefight fingerprints Northflank alerts on the event type, project, and service, so repeated crashes of the same container update one alert and increment its fire count instead of creating new alerts. See [Grouping & deduplication](https://firefight.app/docs/alerts/grouping-and-deduplication.md). Northflank events carry no severity. When a rule creates an incident from a Northflank alert, set the severity on the rule's outcome, otherwise the workspace default severity applies. :::tip Give the source's routing its own rules via **Settings → Alert Sources → Routing** on the source row. A common setup creates an incident for `event` `is one of` `container:crash` and drops or notifies on everything else. Dry-run it first with [Testing your routing](https://firefight.app/docs/alerts/testing-routing.md). ::: --- Source: https://firefight.app/docs/alerts/routing-rules # Routing rules Routing rules decide what happens to every alert Firefight receives. Each rule pairs a set of conditions with an outcome. Rules run in order, top to bottom, and the first rule whose conditions all match wins. Nothing after it is evaluated. ## Where rules live **Settings → Alert Routing** holds the workspace default routing, the fallback for every source. Each alert source can also carry its own rules, reachable from the **Routing** link on the source row in **Settings → Alert Sources**. At ingest, a source's own routing wins when it exists and is enabled, otherwise the workspace default applies. From the routing page you can add rules, edit them, reorder them with the up and down arrows, toggle individual rules on and off, and toggle the whole policy. Disabled rules are skipped during evaluation. If no rule matches, the alert is stored as unmatched. It shows up in **Settings → Alerts** with an unmatched routing state, and nothing is created or sent. ## Conditions A condition names an alert field, an operator, and usually a value. A rule matches only when all of its conditions match. A rule with no conditions always matches, which makes it a natural catch-all at the bottom of the list. | Operator | Matches when | Value | |---|---|---| | is one of | The field equals any value in the list | A list of values | | contains | The field contains the value as a substring | A single value | | starts with | The field begins with the value | A single value | | matches regex | The field matches the regular expression | A regex pattern | | is empty | The field is missing or blank | None | Invalid regex patterns are rejected when you save the rule. Conditions match against the alert's normalized fields, such as `service`, `title`, `severity_raw`, `team`, `environment`, and for Northflank `event`. Two extra fields are always available, `source` carries the alert source's name and `provider` carries its type (`generic` or `northflank`). When the alert's `service`, `team`, `environment`, or `functionality` field names an entry in your catalog, Firefight enriches the alert before evaluation. Related entries one hop away are merged in, so an alert carrying only `service: checkout` also gets `team` filled from the service's owning team. The entry's attributes become dotted fields like `service.tier`, so a condition such as `service.tier` `is one of` `Critical` routes by catalog metadata the alert itself never sent. When you write an `is one of` condition on a catalog field, the value picker offers your catalog entries directly. ## Outcomes Every rule ends in one of four actions. | Action | What it does | |---|---| | Create incident | Declares an incident named after the alert, with the alert's description as the summary. Grouping applies first, so a matching open alert group attaches instead of creating a duplicate. | | Attach to open incident | Attaches the alert to the incident of an open alert group with matching grouping fields. If no such incident is open, nothing is created. | | Notify only | Posts an alert message to a target without creating an incident. | | Drop | Stores the alert and does nothing else. Useful for known noise. | ## Setting severity For the two incident actions you choose the severity of the incident the rule creates. Pick a specific severity, or pick **Severity from source map** to use the source's severity mapping on the alert's raw severity value, falling back to the workspace default severity when unmapped. ## Notify targets A **Notify only** rule sends its message to one of three targets. | Target | Where the message goes | |---|---| | Channel | A Slack channel you pick | | Person (DM) | A direct message to a workspace member | | Owning team's channel | Resolved from the catalog at fire time. The service's own channel wins when set, otherwise the owning team's channel | The owning team target follows your catalog, so reorganizing ownership there redirects alerts without touching any rule. ## Invite targets Rules with an incident action can invite people into the incident channel. You can toggle **Owning team**, which resolves the alert's service to its owning team in the catalog and invites the team's members and manager, and you can add specific people by name. Resolution happens at fire time, and anything that cannot be resolved is noted on the incident rather than blocking its creation. :::note Owning-team resolution depends on your catalog. The service entry needs a relationship to a team entry, and the team type needs attributes marked with the Members, Manager, or Notification channel [routing roles](https://firefight.app/docs/catalog/entries-and-attributes#routing-roles), which the built-in types have out of the box. When something is missing, the incident is still created and the gap is recorded as a note, and this page shows a warning when a role your rules depend on is not set on any attribute. ::: ## Ordering matters Because the first match wins, put your most specific rules at the top and broad catch-alls at the bottom. A wide rule placed early can shadow everything under it. The per-rule test in **Settings → Alert Routing** warns you when a rule's own sample alert is captured by an earlier rule. See [Testing your routing](https://firefight.app/docs/alerts/testing-routing.md). Rules that create incidents interact with the grouping window, which is configured on the same page. See [Grouping & deduplication](https://firefight.app/docs/alerts/grouping-and-deduplication.md). ## Over the API and MCP Rules can be listed, created, changed and removed at `/api/v1/routing_rules`, and with the `upsert_routing_rule` and `delete_routing_rule` MCP tools. A rule is addressed by its priority within its scope, which is either the workspace or one alert source named with `source`. Deleting one leaves a gap rather than renumbering, so read the list back before addressing another rule by priority. See [Using the API](https://firefight.app/docs/api/using-the-api#alert-routing-rules). --- Source: https://firefight.app/docs/alerts/testing-routing # Testing your routing Routing mistakes are cheap to catch before real alerts flow and expensive after. Firefight gives you dry-run tools in **Settings → Alert Routing** and a matching tool for connected AI agents. Dry runs are pure evaluation. No incident is created, no alert is stored, and no message is sent. ## Test a single rule Each rule row in **Settings → Alert Routing** has a test button. Firefight builds a sample alert from the rule's own conditions, for example a `service` `is one of` `checkout` condition produces a sample with `service` set to `checkout`, and runs it through the full first-match evaluation. The result badge on the rule shows whether the sample created an incident, attached, notified, or dropped. Crucially, the test runs against all rules in order, not just the one you clicked. If the sample is captured by an earlier rule, the row warns you that the rule is shadowed, which is the most common routing bug. Rules with a **matches regex** condition cannot generate their own sample, because a pattern does not imply a concrete value. Test those with a custom alert instead. ## Test a custom alert The **Test custom alert** button opens a dialog where you enter the fields a real alert would carry as key and value pairs, such as `service`, `title`, or `severity_raw`. **Run test** evaluates them and shows: - Which rule matched, by its position in the list, or that nothing matched. - The winning action, such as **Create incident** or **Notify only**. - Who would be notified and who would be invited, resolved against your catalog and members the same way a real alert resolves them. - Warnings for anything that would not resolve, such as a team with no channel set. When testing a source's own routing, the evaluation uses that source's rules with the workspace default as fallback, exactly like real ingest. ## Send a test message When a custom test matches a rule with a notify target, the result offers **Send test message**. This is the one testing action with a real side effect. It posts a single clearly labeled message to the resolved target so you can confirm the destination is right and Firefight can post there. It still creates no alert and no incident. ## Dry runs from a connected AI agent `POST /api/v1/routing/evaluate` does the same evaluation over the [REST API](https://firefight.app/docs/api/using-the-api#alert-routing-rules), and agents connected over MCP get an `evaluate_routing` tool for it. The agent passes hypothetical alert fields and optionally an alert source name, and gets back whether a rule matched, the matched rule's position, the outcome with notify and invite targets resolved to names, the enriched field context, and a per-condition trace showing exactly why each rule did or did not match. This lets you ask an agent questions like "what would happen if the checkout service fired a critical alert right now" and get a grounded answer from your live rules. The tool is read-only and changes nothing. ## A pre-launch checklist 1. Run each rule's per-rule test and fix any shadowing warnings. 2. Run custom tests for your most important services and confirm the matched rule, severity, and targets. 3. Use **Send test message** once per notify target to verify delivery. 4. Send one real request from your monitoring tool and check it appears in **Settings → Alerts** with the routing state you expect. :::note Dry runs evaluate rules, they do not exercise ingest. Authentication, field mapping, and batch fan-out only run on a real request to the source URL, so finish with step 4 before trusting the pipeline end to end. See [The generic webhook source](https://firefight.app/docs/alerts/generic-webhook.md). ::: Related pages: [Routing rules](https://firefight.app/docs/alerts/routing-rules.md), [How alerts work](https://firefight.app/docs/alerts/how-alerts-work.md). --- Source: https://firefight.app/docs/api/integrations # Connect an integration An integration is a connection to a tool you already use, so Firefight and the AI agents you connect can read from it and act in it during an incident. Connecting one never turns anything on by itself. Firefight discovers what the tool can do, and you pick which of those capabilities to allow. Go to **Configure → Integrations** to see the catalog. Connecting and managing integrations requires admin access. ## What you can connect Firefight lists these out of the box, grouped by what they are for. | Category | Integrations | |---|---| | Code | GitHub, GitLab | | Issues | Linear | | Telemetry | New Relic, Datadog, Grafana | | Errors | Sentry | | Knowledge | Notion, Confluence | | Databases | PlanetScale, Neon, Supabase, PostgreSQL | | Custom | Any MCP server | If the tool you want is not listed but speaks [Model Context Protocol](https://modelcontextprotocol.io), connect it with **Custom MCP server** and point Firefight at its URL. Nothing about the rest of this page changes. ## Connecting Click **Connect** on the integration you want. Most connect in one click. ### One click GitHub, GitLab, Datadog, Sentry, Linear, Notion, Neon, Supabase and PlanetScale all host their own servers, so you click **Continue with**, approve access on that tool's own consent screen, and land back in Firefight connected. There are no keys to copy and nothing to paste. The consent screen belongs to the tool, not to Firefight, so it is the right place to narrow access. Several tools let you choose exactly what to share there. - GitHub asks you to install its app and pick which repositories it covers - PlanetScale asks you to tick which organizations and databases to grant :::caution Whatever you leave unticked on that screen, Firefight cannot reach later. If a capability keeps failing with a permissions error from the tool, reconnect and check what you granted there. ::: ### With a token New Relic, Grafana, Confluence and PostgreSQL connect with a token instead. Choose **Use a token instead**, paste the server URL and an authorization header, and use a read-only key where the tool offers one. Tokens are stored encrypted, scoped to that one connection, and never shown again after you save them. They are never sent to agents. ## Choosing what an integration can do Once connected, open the connection to see every capability the tool offers. All of them arrive switched off. Switch on only what you want available. Each capability you enable becomes a permission you can then grant to people and agents, described in [Permissions](https://firefight.app/docs/gateway/permissions.md). For long lists, use the controls above the capabilities. | Control | What it does | |---|---| | Reads only | Turns on every capability that only reads, and turns off every one that writes | | Enable all | Turns on everything the tool offers, including capabilities that write | | Disable all | Turns everything back off | Capabilities that change something in the connected tool are marked **write**. The header tells you how many are on, and how many of those write, so widening access is visible at a glance. :::tip **Reads only** is the safer starting point when you are connecting a tool for AI agents to investigate with. You can always switch individual write capabilities on afterwards. ::: ## Environments If your workspace has environments in the [service catalog](https://firefight.app/docs/catalog/overview.md), a connection can hold separate credentials for each one, so production and development are reached with different access. Connect the same integration twice, choosing a different environment each time. You get one connection with two sets of credentials, one tool list, and the same capabilities. What changes is who is allowed to use it where, which you control in [Permissions](https://firefight.app/docs/gateway/permissions.md). Connecting the same tool twice for the same environment is a different thing. Use it when you have two separate accounts, such as two GitHub organizations or two PlanetScale organizations. Choose **Add a second account** and give it a name. Its capabilities and permissions stay entirely separate from the first. To change which environment an existing connection's credentials answer for, open the connection and use the dropdown next to those credentials. ## Keeping a connection healthy Each set of credentials shows its own status. | Status | Meaning | |---|---| | Healthy | Firefight reached the tool and read its capability list | | Failing | Firefight could not reach the tool, with the error shown on the connection | | Not checked | Firefight has not contacted the tool yet | **Refresh tools** re-reads the capability list and rechecks health. Run it after the tool adds features you want, or when you are diagnosing a failure. If a capability disappears from the tool, it switches off and stops being available. Should it come back, the permissions you granted for it are still there, so you do not have to set them up again. ## Turning an integration off The toggle on an integration is a kill switch. Turning it off immediately removes its capabilities from everything, including agents connected over MCP, without losing your configuration. Turn it back on and everything returns. **Disconnect** removes the connection from your list and stops every capability it provided. :::caution Disconnecting stops Firefight using the access, but it does not revoke it at the other end. To cut access completely, also revoke Firefight in that tool's own settings. ::: --- Source: https://firefight.app/docs/api/mcp-server # Connect AI agents (MCP) Firefight ships a [Model Context Protocol](https://modelcontextprotocol.io) server, so AI agents like Claude can take part in your incidents directly. An agent connected over MCP can search incidents, pull a full incident timeline, inspect alerts and how they routed, look up ownership in the service catalog, and dry-run your alert routing rules. With permission it can also work an incident the way a person does: declare one, raise and pick up work, claim a runbook step, pull people into the channel, ask a named person to respond, and close it out. It can maintain your catalog, runbooks and routing rules, and resolve approvals, too. Everything an agent does is recorded under **Gateway → Activity**. ## Endpoint The server speaks Streamable HTTP at a single endpoint. ``` POST https://app.firefight.app/mcp ``` Every request is self-contained, and all data is scoped to the workspace of the token you authenticate with. ## Authenticating There are three ways to connect, depending on what is doing the connecting. ### OAuth, for interactive clients Clients that support MCP OAuth, such as Claude Code, connect without you ever handling a token. Add the server and start using it. ```sh claude mcp add --transport http firefight https://app.firefight.app/mcp ``` On first use, your browser opens Firefight's consent screen, which names the client asking for access and the workspace it would reach. If you belong to more than one workspace, pick the one you want it to reach. Click Authorize and the client is connected. The client registers itself automatically through dynamic client registration, PKCE is required, and access tokens are short-lived with refresh rotation. The connection acts as you, with the same access you have in the workspace. ### Agent token, for an AI acting as itself An AI that takes part in incidents gets its own identity rather than borrowing yours. Create it under **Gateway → Agents**, which hands you its token once, and configure your client to send that. ```sh claude mcp add --transport http firefight https://app.firefight.app/mcp \ --header "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" ``` Everything the agent does is recorded under its own name, so an incident it declared says the agent declared it, and an action item it picked up shows the agent holding it. It holds only the abilities you granted it, whatever access you have yourself. Agent tokens do not expire, because what a leaked one can do is bounded by its grants and its approval rules rather than by how soon it has to be renewed. Rotate one at any time from its row. See [Agents](https://firefight.app/docs/gateway/agents.md). ### API key, for pipelines and CI Automation with nothing to say on a timeline passes a Firefight API key as a Bearer token instead. Create one under **Settings → API Keys**, then configure your client to send it. ```sh claude mcp add --transport http firefight https://app.firefight.app/mcp \ --header "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" ``` For Cursor, put the same thing in `.cursor/mcp.json`. ```json { "mcpServers": { "firefight": { "url": "https://app.firefight.app/mcp", "headers": { "Authorization": "Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" } } } } ``` Any other MCP client works the same way, either through OAuth discovery at `/.well-known/oauth-protected-resource` or with the `Authorization: Bearer` header. ## Multiple workspaces A connection reaches one workspace, whichever way you authenticated. An OAuth connection is granted for the single workspace you picked on the consent screen, and an API key belongs to the workspace you created it in. There is no setting that widens a connection to cover more. To give an agent a second workspace, add the server again under a different name. ```sh claude mcp add --transport http firefight-acme https://app.firefight.app/mcp ``` Authorize that one and pick the second workspace on the consent screen. For a headless agent, create an API key in the second workspace and pass it the same way. Each connection stands on its own, so its tools appear under the name you gave it, and revoking one leaves the other working. An agent is told which workspace it reaches as soon as it connects, so you can ask it directly rather than working it out from what it returns. The **Connected agents** list under **Settings → API Keys** shows the connections for the workspace you are viewing. Switch workspaces to see the rest. ## Tools These tools read your workspace and change nothing. | Tool | What it returns | |---|---| | `search_incidents` | Incident summaries, newest first. Filter by status, severity, lifecycle stage, free-text query, and declared-at time range. | | `get_incident` | One incident in full, by ID or identifier like `INC-42`. Details, timeline events, postmortem status, attached alerts, every incident role with whoever holds it, the open action items, and the attached runbooks with their steps. | | `search_alerts` | Ingested alerts, newest first. Each shows its source, firing status, routing state, the rule that matched, and the incident it attached to. | | `search_catalog` | Catalog entries with their attributes and relationships, such as which team owns a service. Filter by type, name query, or exact slug. | | `evaluate_routing` | A dry run of alert routing against hypothetical alert fields. Returns the matching rule, the outcome, a per-condition trace, and warnings when a [routing role](https://firefight.app/docs/catalog/entries-and-attributes#routing-roles) your rules depend on is not set on any catalog attribute. Nothing is created or notified. | | `search_runbooks` | Runbook listings with name, summary, and step count. Filter by a name or summary query. | | `get_runbook` | One runbook in full, by slug. The complete procedure content, ordered steps, and external link. | | `search_approvals` | Approval requests and their status. | | `list_abilities` | Everything that can be granted, with its risk level and whether an approval rule can hold it. | | `list_principals` | The people, agents and service keys of the workspace, each with the grants it holds. | | `search_activity` | The gateway's audit log, filtered by decision or ability. | | `get_form` | One lifecycle form and every field on it, hidden ones included, with each field's type, whether it is visible, required, or locked, and any conditions gating it. | | `get_workspace_config` | How the workspace is set up, in one call: severities, statuses with their lifecycle stage, incident types, incident roles, alert sources and webhooks. Disabled entries are included and marked. | | `list_agents` | The agents in the workspace, each with its live tokens and how many abilities it holds. | | `list_api_keys` | The workspace's service keys, with the abilities each holds and when it was last used. Never the token itself. | | `get_postmortem` | The postmortem written for an incident, as HTML, with its status and who wrote it. | | `get_incident_transcript` | What people said in the incident's channel, oldest last, with who said it and when. Needs its own ability and the workspace setting. | ### Tools that change things These need a credential with permission to make the change, and every call is recorded. | Tool | What it does | |---|---| | `upsert_catalog_entry` | Creates or updates a catalog entry | | `delete_catalog_entry` | Removes a catalog entry | | `upsert_catalog_type` | Creates or changes a kind of thing the catalog holds, and the attributes its entries carry | | `delete_catalog_type` | Removes a catalog type and the entries under it | | `upsert_routing_rule` | Creates or updates an alert routing rule | | `delete_routing_rule` | Removes a routing rule | | `update_routing_config` | Changes routing configuration | | `upsert_runbook` | Creates or updates a runbook | | `upsert_custom_field` | Creates or updates a custom field and its options | | `upsert_form_field` | Attaches a field to a form, or changes its visibility, required setting, or conditions | | `upsert_severity` / `delete_severity` | Creates, changes or removes a severity | | `upsert_status` / `delete_status` | Creates, changes or removes a status | | `upsert_incident_type` / `delete_incident_type` | Creates, changes or removes an incident type | | `upsert_incident_role` / `delete_incident_role` | Creates, changes or removes an incident role | | `upsert_alert_source` / `delete_alert_source` | Creates, changes or removes an alert source | | `upsert_webhook` / `delete_webhook` | Creates, changes or removes an outbound webhook | | `test_webhook` | Sends a test delivery to an outbound webhook, replaying the newest matching event from your workspace | | `upsert_agent` / `delete_agent` | Creates, changes or removes an agent | | `rotate_agent_token` / `revoke_agent_token` | Issues a new token for an agent, or ends one | | `upsert_api_key` / `delete_api_key` | Creates, changes or removes a service key | | `declare_incident` | Opens a new incident and its channel | | `post_incident_update` | Posts an update on an incident | | `resolve_incident` | Closes an incident | | `cancel_incident` | Cancels an incident that turned out not to be one | | `reopen_incident` | Reopens an incident that came back | | `create_action_item` | Adds a piece of work to an incident and posts it in the channel | | `assign_action_item` | Takes a piece of work, or hands it to someone else | | `complete_action_item` | Marks a piece of work done | | `claim_runbook_step` | Takes one step of an attached runbook | | `escalate_incident` | Asks a named person to pick the incident up | | `invite_responders` | Brings people into the incident channel | | `link_incident` | Links two incidents, or marks one a duplicate of the other | | `give_shoutout` | Thanks someone for their work on the incident | | `assign_incident_role` | Puts one person in an incident role, or clears it | | `attach_runbook` | Attaches a runbook to an incident | | `dismiss_timeline_note` | Removes one AI note from an incident's timeline | | `approve_approval` | Approves a pending request | | `deny_approval` | Denies a pending request | | `upsert_permission_set` | Creates or updates a permission set | | `delete_permission_set` | Removes a permission set from everyone holding it | | `grant_ability` | Grants an ability or a set to a person, agent or key | | `revoke_grant` | Revokes a grant | | `upsert_approval_rule` | Creates or updates an approval rule | | `delete_approval_rule` | Removes an approval rule | | `start_postmortem` | Opens the postmortem for a resolved incident, empty or drafted by AI | | `update_postmortem` | Replaces the body of a postmortem | | `set_postmortem_status` | Moves a postmortem to draft, in progress, in review or completed | `upsert_catalog_entry` takes attributes by name, and an attribute pointing at another entry accepts that entry's slug, so an agent can set which team owns a service without looking up an ID. `upsert_catalog_type` shapes the catalog itself, adding a kind of thing such as Datastore and saying what every entry under it carries. Attributes are matched by name, so resending a list renames rather than replaces and the entries already holding a value keep it, and sending no attributes leaves the shape alone. A reference attribute names the type it points at by slug. Built-in types keep their slug and their own attributes, and `delete_catalog_type` refuses a type another type points at, naming the attribute in the way. See [Entries, attributes & relationships](https://firefight.app/docs/catalog/entries-and-attributes.md). `upsert_custom_field` and `upsert_form_field` let an agent change what responders are asked during an incident, so they are worth granting deliberately. Options are matched by label, which means resending a list renames rather than replaces, and the incidents already holding an option keep pointing at it. Conditions accept slugs, so you can gate a field on `production` without looking up an ID first. A condition on a custom field names that field by its key and takes option labels or catalog entry slugs as its values, so gating on an Affected service field being `checkout` needs no IDs either. Severity and Status refuse to be hidden or made optional here exactly as they do in the dashboard. Call `get_form` first to see what a form holds, including hidden fields, because an update replaces the set rather than adding to it. The gateway tools let an agent administer access the same way an admin does on the dashboard: see who holds what, bundle abilities into sets, grant and revoke, and decide which abilities wait for approval. They check `permissions`, which only an admin holds and which cannot be granted to a key, so they work over a personal token or an OAuth connection made by an admin, never a service key. `approve_approval` and `deny_approval` decide as whoever is behind the token. Connected as you, an agent can be your chief of staff, spotting a request that names you in `search_approvals` and approving it once you say so. Connected with its own service key, it can decide only the rules that name that key and have **Agents may decide this rule** switched on, and the decision is recorded under the agent's name. See [Permissions](https://firefight.app/docs/gateway/permissions.md) and [Approvals](https://firefight.app/docs/gateway/approvals.md). `attach_runbook` takes the incident and a runbook slug, and posts that runbook's steps in the incident channel. Runbooks whose conditions match attach on their own, so this is for the ones that do not. Attaching the same runbook twice does nothing the second time. It needs `incidents: update`, not a runbooks permission, because it changes the incident rather than the runbook. `dismiss_timeline_note` takes the incident and the id of a note, and removes that note from the timeline. It is for the case where a note reads a conversation wrong, crediting the wrong person or taking a joke for a decision. The note is kept and marked dismissed rather than deleted, and stops coming back from `get_incident`. Note ids come from that same timeline, where each note also names its kind, so an agent can ask for the root cause of a past incident directly. See [Timeline notes](https://firefight.app/docs/incidents/concepts#timeline-notes). ### Working an incident The incident tools are the same operations a person has in Slack, so an agent using them is a responder rather than a reporter. What it does shows up in the incident channel as it happens, and on the incident's timeline under the agent's own name. `declare_incident`, `post_incident_update`, `resolve_incident`, `cancel_incident` and `reopen_incident` all ask exactly what your workspace's [incident forms](https://firefight.app/docs/customization/incident-forms.md) ask a person, custom fields included. Call `get_form` first with the form you are about to submit, then pass the answers keyed the way it named them. A field the form does not ask for is refused rather than quietly dropped, and a missing required field comes back naming what is missing. `create_action_item` adds a piece of work and posts it in the channel, exactly like `/ff action`. Pass `kind` as `action` for work during the incident or `followup` for work after it, and `member` to hand it straight to someone. `assign_action_item` takes a piece of work when you leave `member` out, which is the **I can take this** button, and hands it over when you name someone else, which announces the handover in the channel. `complete_action_item` marks it done and posts the completion. Both take the item's id, which comes from `get_incident`. `claim_runbook_step` takes one step of a runbook already attached to the incident, which creates the action item behind that step, or hands over the one that already exists. Runbook and step ids come from `get_incident`. `invite_responders` brings people into the incident channel so they can see what is happening. `escalate_incident` is the stronger one: it asks one named person to pick the incident up, posts the ask in the channel, messages them directly with an acknowledge button, and reminds them if they do not answer. Reach for it when the agent has gone as far as it can on its own. Escalating, inviting and giving a shoutout all post in the incident channel, so all three refuse an incident that has been resolved or canceled, and one whose channel Firefight is still creating. The refusal names which it is. `link_incident` records a `related` link on both incidents' timelines and changes nothing else. Passing `duplicate` instead also cancels this incident, naming the one that absorbed it, so use it only when the two really are the same event seen twice. `give_shoutout` thanks a named person in the incident channel, the same as `/ff shoutout`. Every one of these needs `incidents: create` or `incidents: update`, so an agent granted only `incidents: read` can follow an incident and change nothing about it. `assign_incident_role` takes the incident, the role slug, and the person as an email address or Slack user ID. Leave the person out to clear the role. Each role holds one person, so assigning hands the role over from whoever had it. Call `get_incident` first to see which roles exist and who currently holds them. Every role change announces itself in the incident channel, so the tool refuses any of them on a resolved or canceled incident and names the role it would have changed. See [Incident roles](https://firefight.app/docs/incidents/incident-roles.md). ### Writing the postmortem An incident that has been resolved can have its write-up started, edited and moved along without anyone opening the dashboard. `start_postmortem` opens it. Leave `generate` out for an empty one you write yourself, or pass it to have Firefight draft the first version from the incident. Generating takes a moment, so the response comes back immediately with a generation state to poll on. Read it back with `get_postmortem` until the state clears. `update_postmortem` sends the whole document as HTML and replaces what was there rather than appending to it, so read the current one first if you mean to add a section. Every version is kept, so a rewrite can be compared and rolled back from the dashboard. `set_postmortem_status` moves it to `draft`, `in_progress`, `in_review` or `completed`, which is how your team knows whether it still needs writing or reading. `get_postmortem` gives you a `version`, and `update_postmortem` takes it back. This matters because a person is often writing the same document in the dashboard at the same time. If they saved while the agent was composing, the agent's write is refused rather than replacing what they wrote, and the agent reads again and reapplies. Moving the status needs no version, since it changes nothing anybody is typing. Starting one is refused while the incident is still open, and refused outright for a canceled incident, which has nothing to write up. An incident that already has a postmortem is refused too, so read before you start. All four tools check `incidents`, not a permission of their own. See [Postmortems](https://firefight.app/docs/postmortems/overview.md). ### Reading the conversation `get_incident_transcript` returns what people actually said in the incident channel, which is where the reasoning lives that a timeline only summarises. Read it when you know what happened and need to know why. It needs the **Incident Transcripts** ability, which is separate from `incidents` on purpose, and the workspace has to have turned transcript access on under **Settings → Workspace**. An agent granted every incident ability reads nothing until both are true. See [Incident conversations](https://firefight.app/docs/workspace/incident-conversations.md). Messages come back oldest last, 100 at a time and 500 at most, with a `more_before` cursor for walking further into the past. Each message says whether anything in it was redacted. ### Setting the workspace up Every settings screen has tools behind it. That means you can ask an agent connected as you to do the setup, rather than clicking through it. "Add a SEV0 above our current top severity, and a Mitigating status in the active stage" is one conversation instead of two screens. Call `get_workspace_config` first. It returns your severities, statuses, incident types, incident roles, alert sources and webhooks in one response, and the slugs it gives back are what every other tool takes. Each list has its own pair of tools rather than one tool with a type argument, because the lists do not ask for the same things. A status belongs to a lifecycle stage, a severity's position says how severe it is, and only some are colored or have a default. Each of these four lists takes `position`, 1 being first, so "add a SEV0 above our current top severity" is `position: 1`. Separate tools mean the fields an agent is offered are the fields that list actually has. Passing a slug changes the one that already has it, and leaving the slug out creates a new one. **Renaming never moves the slug**, because the slug is what your existing incidents point at. A slug that matches nothing comes back as an error rather than quietly creating a second entry under a fresh name. Deleting is refused while anything still points at the entry, and the refusal says how many. Disabling is what you want then. A disabled entry keeps its slug and stops being offered to responders, and `get_workspace_config` still returns it so an agent can turn it back on. :::caution `upsert_form_field`, `upsert_custom_field` and the severity and status tools change what responders are asked and what they can choose during an incident. Grant them deliberately. ::: ### Agents and keys Creating an agent, rotating its token and minting a service key are all available here, and they behave differently from everything else on this page. They check `permissions` and `api_keys`, which are admin-only and cannot be granted to anyone. A connection acting as an admin reaches them. **A service key or an agent never can, whatever it has been granted**, so an agent cannot create another agent or widen its own access. `upsert_agent` returns the new agent's token once, in that response, and never again. So does `rotate_agent_token`. Neither `list_agents` nor `list_api_keys` ever carries a token. Rotating is an overlap rather than a swap. The new token works immediately while the old one keeps working, so the agent stays up while you update its configuration, and `revoke_agent_token` ends the old one when you are ready. See [Agents](https://firefight.app/docs/gateway/agents.md). ### Tools from your integrations Every capability you switch on under **Configure → Integrations** also appears here, so an agent can query PlanetScale or open a Linear issue through the same connection. Nothing appears until you enable it, and who may call it is set in [Permissions](https://firefight.app/docs/gateway/permissions.md). An agent's tool list only carries the capabilities its key may call, so a narrowly scoped key never sees the rest of your connections. See [Connect an integration](https://firefight.app/docs/api/integrations.md). :::tip `evaluate_routing` is a safe way to let an agent answer "what would happen if this alert fired?" It evaluates your real routing rules but never creates incidents or sends notifications. ::: ## Result caps Search tools return 25 results by default and accept a `limit` of up to 50. When more results exist than the cap allows, the response includes a `truncated` flag set to `true`, so agents know to narrow their filters instead of assuming they saw everything. Incident timelines in `get_incident` are capped at 50 events with their own `timeline_truncated` flag. Each token may make up to 1,000 MCP requests per minute. ## Permissions What a connected agent can query depends on the credential it authenticates with. | Credential | Access | |---|---| | OAuth connection | Acts as you, with your access | | Personal token | Acts as you, with your access | | Agent token | Acts as the agent, with only what the agent was granted | | Service key | Only what it has been granted, and nothing by default | Agent tokens and service keys map to tools by resource. | Tool | Required permission | |---|---| | `search_incidents` | `incidents: read` | | `get_incident` | `incidents: read` | | `search_alerts` | `alerts: read` | | `search_catalog` | `catalog: read` | | `evaluate_routing` | `policies: read` | | `search_runbooks` | `runbooks: read` | | `get_runbook` | `runbooks: read` | | `search_approvals` | `approvals: read` | | `get_form` | `forms: read` | | `upsert_catalog_entry` | `catalog: create` or `catalog: update` | | `delete_catalog_entry` | `catalog: delete` | | `upsert_routing_rule` | `policies: create` or `policies: update` | | `delete_routing_rule` | `policies: delete` | | `update_routing_config` | `policies: update` | | `upsert_runbook` | `runbooks: create` or `runbooks: update` | | `upsert_custom_field` | `custom_fields: create` or `custom_fields: update` | | `get_workspace_config` | `incidents: read` | | `upsert_severity` / `delete_severity` | `severities: create`, `update` or `delete` | | `upsert_status` / `delete_status` | `statuses: create`, `update` or `delete` | | `upsert_incident_type` / `delete_incident_type` | `incident_types: create`, `update` or `delete` | | `upsert_incident_role` / `delete_incident_role` | `incident_roles: create`, `update` or `delete` | | `upsert_alert_source` / `delete_alert_source` | `alerts: create`, `update` or `delete` | | `upsert_webhook` / `delete_webhook` | `webhooks: create`, `update` or `delete` | | `test_webhook` | `webhooks: update` | | `list_agents`, `upsert_agent`, `rotate_agent_token`, `revoke_agent_token`, `delete_agent` | `permissions`, admin-only and ungrantable | | `list_api_keys`, `upsert_api_key`, `delete_api_key` | `api_keys`, admin-only and ungrantable | | `upsert_form_field` | `forms: update` | | `declare_incident` | `incidents: create` | | `post_incident_update` | `incidents: update` | | `resolve_incident` | `incidents: update` | | `cancel_incident` | `incidents: update` | | `reopen_incident` | `incidents: update` | | `create_action_item` | `incidents: update` | | `assign_action_item` | `incidents: update` | | `complete_action_item` | `incidents: update` | | `claim_runbook_step` | `incidents: update` | | `escalate_incident` | `incidents: update` | | `invite_responders` | `incidents: update` | | `link_incident` | `incidents: update` | | `give_shoutout` | `incidents: update` | | `assign_incident_role` | `incidents: update` | | `attach_runbook` | `incidents: update` | | `dismiss_timeline_note` | `incidents: update` | | `approve_approval` | `approvals: update` | | `deny_approval` | `approvals: update` | | `list_abilities` | `permissions: read` | | `list_principals` | `permissions: read` | | `search_activity` | `permissions: read` | | `upsert_permission_set` | `permissions: create` or `permissions: update` | | `delete_permission_set` | `permissions: delete` | | `grant_ability` | `permissions: create` | | `revoke_grant` | `permissions: delete` | | `upsert_approval_rule` | `permissions: create` or `permissions: update` | | `delete_approval_rule` | `permissions: delete` | A tool call without the needed permission returns an error naming the missing permission, and the agent can keep using the tools it does have. See [API overview](https://firefight.app/docs/api/overview.md) for how keys and permissions work. ## Managing connected agents Everything connected over OAuth appears under **Settings → API Keys** in the Connected agents section, where you can revoke any connection. Revoking immediately invalidates its tokens, and the client has to go through consent again to reconnect. Automation using an API key is managed like any other key on the same page: deactivate or delete the key to cut it off. An AI with its own identity lives under **Gateway → Agents** instead, where you rotate or revoke its tokens, change what it may do, or switch it off without losing what it did. See [Agents](https://firefight.app/docs/gateway/agents.md). --- Source: https://firefight.app/docs/api/overview # API overview Firefight exposes a JSON REST API for reading and writing incidents, reading your workspace's configuration, and managing service catalog entries. Anything you can automate around your incident response, from declaring incidents out of your monitoring pipeline to syncing the catalog from your infrastructure, goes through this API. ## Base URL and versioning All endpoints live under a versioned path. ``` https://app.firefight.app/api/v1 ``` The version is part of the URL. Breaking changes ship as a new version, so `v1` responses stay stable. ## Machine-readable description Everything on this page and the next is also published as an OpenAPI document, so you can generate a client, load the API into an explorer, or hand it to an AI agent as a tool list. ``` https://firefight.app/openapi.json https://firefight.app/openapi.yaml ``` Every operation in it carries a unique ID, a description, typed parameters, and a response schema. The permission an operation checks is on the operation itself as `x-firefight-permission`, so you can tell before calling it which grant a service key needs. ## Creating API keys Create and manage keys under **Settings → API Keys** in the Firefight web app. Every key belongs to your workspace and can optionally carry an expiry date. The token itself starts with `ff_` and is shown once, at creation time. Store it somewhere safe, because Firefight keeps only a hash. There are two kinds of key. | Kind | Who can create it | What it can do | |---|---|---| | Service key | Workspace admins | A standalone integration identity with granular permissions you pick per resource and action. Use it for monitoring pipelines, CI, and other headless automation. | | Personal token | Any member, for themselves | Acts as you. It reads everything you can see, and writes only what your own role lets you change. Use it for your own scripts and agent sessions. | A personal token carries your own authority and nothing more, so a member's token reads without writing, and an admin's token can also change workspace configuration. For headless automation, mint a service key with exactly the permissions it needs rather than lending it yours. You can rename a key, edit a service key's permissions, deactivate it temporarily, or delete it at any time from the same settings page. ## Authentication Pass your token in the `Authorization` header on every request. The one exception is [sending an alert in](https://firefight.app/docs/alerts/generic-webhook.md), which authenticates with the alert source's own token rather than an API key. ``` Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA ``` Requests without a valid token get a `401` response. Requests with a valid token that lacks the needed permission get a `403`. ## Resources and permissions Service key permissions are a matrix of resources and actions. A key only holds the permissions you grant it, and each endpoint checks one specific pair. | Resource | Actions with effect | What they unlock | |---|---|---| | `incidents` | `read`, `create`, `update` | Read, declare and update incidents, read a timeline and dismiss a note, assign incident roles, and take part: raise and pick up work, claim a runbook step, escalate, invite, link, and give a shoutout | | `severities` | `read`, `create`, `update`, `delete` | List your severity levels, and manage them | | `statuses` | `read`, `create`, `update`, `delete` | List your incident statuses, and manage them | | `incident_types` | `read`, `create`, `update`, `delete` | List your incident types, and manage them | | `custom_fields` | `read`, `create`, `update` | List your custom field definitions, and create or change one | | `forms` | `read`, `update` | Read a lifecycle form, and change which fields it asks for | | `catalog` | `read`, `create`, `update`, `delete` | Read catalog types, and manage catalog entries | | `runbooks` | `read`, `create`, `update` | List and fetch runbooks, and create or change one | | `alerts` | `read`, `create`, `update`, `delete` | Search ingested alerts, and manage alert sources | | `policies` | `read`, `create`, `update`, `delete` | Dry-run alert routing, and manage routing rules | | `approvals` | `read`, `update` | List approval requests, and approve or deny one | | `permissions` | admins only | Abilities, principals, permission sets, grants, approval rules and the activity log. Cannot be granted to a key, so use a personal token | | `incident_roles` | `read`, `create`, `update`, `delete` | List your incident roles, and manage them | | `webhooks` | `read`, `create`, `update`, `delete` | List outbound webhooks, and manage them | | `api_keys` | admins only | Service keys. Cannot be granted to a key, so use a personal token | Almost every pair has a REST endpoint behind it. Creating and changing runbooks, and managing alert routing rules, are the two exceptions, and both are [MCP](https://firefight.app/docs/api/mcp-server.md) tools instead. Everything else on this table is reachable over REST, over MCP, and on the dashboard. The same resources are what you grant a person under [Permissions](https://firefight.app/docs/gateway/permissions.md). **Which surface a request came through makes no difference to what it is allowed to do**, because the permission is checked the same way in all three. A key that can change severities on the dashboard can change them over the API and over MCP, and one that cannot change them anywhere cannot change them here. See [Configuring your workspace](https://firefight.app/docs/api/using-the-api#configuring-your-workspace). `custom_fields` and `forms` are deliberately separate. `custom_fields` covers what a field is, such as adding an Affected region dropdown. `forms` covers which fields responders are asked for and when, so a key that lets an agent define a field does not also let it make Name required on every declaration. See [Incident forms](https://firefight.app/docs/customization/incident-forms.md). :::tip Grant the narrowest set that works. A key that only declares incidents from your alerting pipeline needs `incidents:create` and nothing else. ::: ## Rate limits Each key may make up to 1,000 requests per minute. Beyond that, requests get a `429` response until the window resets. See [Using the API](https://firefight.app/docs/api/using-the-api.md) for the error format. ## Your first request List the most recent incidents in your workspace. ```bash curl https://app.firefight.app/api/v1/incidents \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" ``` ```json { "incidents": [ { "id": "0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70", "identifier": "INC-42", "name": "Checkout latency spike in eu-west", "summary": "p99 latency on the checkout service exceeded 4s.", "status": { "id": "b1f6a2d4-8c3e-4f5a-9b7d-1e2c3a4b5d6f", "name": "Investigating", "lifecycle_stage": "active" }, "severity": { "id": "c7e2f9a1-3b4d-4c6e-8f0a-5d6b7c8e9f01", "name": "SEV2", "rank": 2 }, "type": null, "lead": { "id": "e4a8b2c6-1d3f-4e5a-b7c9-0f1a2b3c4d5e", "type": "user", "name": "Maja Kovac", "email": "maja@example.com" }, "source": "datadog", "declared_by": null, "declared_at": "2026-07-18T09:12:44Z", "detected_at": null, "resolved_at": null, "created_at": "2026-07-18T09:12:44Z", "updated_at": "2026-07-18T09:40:02Z", "custom_fields": {}, "visibility": "public" } ], "pagination": { "page": 1, "per_page": 25, "total": 1, "total_pages": 1 } } ``` From here, read [Using the API](https://firefight.app/docs/api/using-the-api.md) for idempotent incident creation, error handling, and pagination, or set up [outbound webhooks](https://firefight.app/docs/api/webhooks.md) to push events to your systems instead of polling. --- Source: https://firefight.app/docs/api/using-the-api # Using the API This page covers the mechanics you need for reliable integrations. It assumes you already have a key from [API overview](https://firefight.app/docs/api/overview.md). ## Idempotent incident creation `POST /api/v1/incidents` requires an `idempotency_key` in the request body. It is any string you choose that uniquely identifies this creation attempt, for example your alert's ID or a UUID you generate. The key protects you from duplicates when a request times out and you retry. | Situation | Response | |---|---| | First request with a key | `201 Created`, a new incident | | Repeated request with the same key | `200 OK`, the incident created the first time | Idempotency keys expire after 24 hours. After that, reusing the same key creates a new incident, so treat keys as protection for retries, not as a long-term deduplication mechanism. :::tip Derive the key from the triggering event, such as `pagerduty-alert-8842`. Retries of the same event then dedupe naturally, no matter which of your workers sends them. ::: ## Errors Every error is JSON with a single `error` object. The `request_id` identifies the request, so include it when you contact support. ```json { "error": { "type": "validation_error", "message": "Validation failed: Name can't be blank", "request_id": "9f2c1d4e-8a7b-4c3d-b5e6-0a1b2c3d4e5f", "errors": [ { "field": "name", "message": "can't be blank" } ] } } ``` The `errors` array appears only on validation errors and lists each failing field. | Status | `type` | When | |---|---|---| | 400 | `bad_request` | A required parameter is missing | | 401 | `unauthorized` | The token is missing, invalid, expired, or deactivated | | 403 | `forbidden` | The token lacks the permission this endpoint checks | | 404 | `not_found` | The resource does not exist in your workspace | | 422 | `validation_error` | The request was well-formed but the data is invalid | | 422 | `incident_not_active` | The incident is resolved or canceled and cannot take this change | | 429 | `rate_limit_exceeded` | You exceeded 1,000 requests per minute for this key | ## Pagination List endpoints that can grow large, such as incidents and catalog entries, accept two query parameters. | Parameter | Default | Maximum | |---|---|---| | `page` | 1 | none | | `per_page` | 25 | 100 | Every paginated response includes a `pagination` object. ```json { "pagination": { "page": 2, "per_page": 25, "total": 132, "total_pages": 6 } } ``` The incident list also accepts filters as query parameters. Use `severity_id`, `status_id`, or `lifecycle_stage` with one of `triage`, `active`, `closed`, or `canceled`. Look up the IDs from `GET /api/v1/severities` and `GET /api/v1/statuses`. ## Reading an incident's timeline `GET /api/v1/incidents/:id/timeline` returns everything recorded against an incident in order, paginated like any other list. Each event carries its type, a plain-language description, who did it, and when. Events that Firefight noted from the channel also carry a `milestone` object. That is how you read how an incident was debugged without touching Slack. ```json { "event_type": "milestone.noted", "description": "noted Diego identified the migration lock as the root cause", "automated": true, "occurred_at": "2026-08-25T14:31:00Z", "actor": null, "milestone": { "kind": "root_cause", "statement": "Diego identified the migration lock as the root cause", "said_by": "Diego", "said_at": "2026-08-25T14:31:00Z", "message_text": "found it, the migration is holding the lock", "permalink": "https://yourteam.slack.com/archives/C123/p1756132260", "dismissed_at": null } } ``` Every other event has `milestone: null`. The `kind` is one of `hypothesis`, `finding`, `root_cause`, `mitigation`, `decision`, `blocker`, `impact`, or `recovery`. To remove a note that reads a conversation wrong, `PATCH /api/v1/incidents/:id/timeline/:note_id/dismiss` with an `incidents: update` key. It returns the dismissed note, and that note stops appearing in the timeline afterwards. Anything that is not a note is refused with a `422`. See [Timeline notes](https://firefight.app/docs/incidents/concepts#timeline-notes). ## Taking part in an incident Changing an incident's status is `PATCH /api/v1/incidents/:id`. Everything else a responder does inside an incident has its own endpoint, and each one needs `incidents: update`. | Endpoint | What it does | |---|---| | `GET /api/v1/incidents/:id/action_items` | Lists the open work, with who holds each piece | | `POST /api/v1/incidents/:id/action_items` | Adds a piece of work and posts it in the channel | | `PATCH /api/v1/incidents/:id/action_items/:item_id` | Takes it, hands it over, or marks it done | | `POST /api/v1/incidents/:id/escalate` | Asks a named person to pick the incident up | | `POST /api/v1/incidents/:id/invite` | Brings people into the incident channel | | `POST /api/v1/incidents/:id/link` | Links two incidents, or marks one a duplicate | | `POST /api/v1/incidents/:id/shoutout` | Thanks someone in the channel | | `POST /api/v1/incidents/:id/runbook_steps/claim` | Takes one step of an attached runbook | Anywhere you name a person, pass their email address or their Slack user ID rather than a Firefight ID. A name that matches nobody in the workspace returns a `404`. Creating a piece of work takes a `description`, an optional `kind` of `action` or `followup`, and an optional `assignee_id` to hand it straight to someone. ```bash curl -X POST https://app.firefight.app/api/v1/incidents/0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70/action_items \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "description": "Drain replica 2", "kind": "action" }' ``` One `PATCH` covers taking a piece of work, handing it over and finishing it, because each is the same sentence: this item now looks like this. Send `assignee_id: null` to take it yourself, name someone else to hand it over, and send `status: "done"` to finish it. ```bash curl -X PATCH https://app.firefight.app/api/v1/incidents/0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70/action_items/6b1f0a29-4c7d-4e1a-9f2b-3d8c5e7a0b14 \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "status": "done" }' ``` Escalating asks one named person to respond, posts the ask in the channel, messages them directly with an acknowledge button, and reminds them if they do not answer. Inviting is the quieter one, it just brings people in. ```bash curl -X POST https://app.firefight.app/api/v1/incidents/0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70/escalate \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "member_id": "priya@example.com", "reason": "Needs a database owner" }' ``` Linking takes `other_incident_id` and a `relationship` of `related` or `duplicate`. `related` records the link on both timelines and changes nothing else. `duplicate` also cancels this incident, naming the one that absorbed it, so a resolved or canceled incident has to be reopened first. ## Configuring your workspace Every settings screen has a matching set of endpoints, so a script can set a workspace up the same way a person does by clicking. | Resource | Endpoints | Addressed by | |---|---|---| | Severities | `/api/v1/severities` | slug | | Statuses | `/api/v1/statuses` | slug | | Incident types | `/api/v1/incident_types` | slug | | Incident roles | `/api/v1/incident_roles` | slug | | Custom fields | `/api/v1/custom_fields` | slug | | Incident forms | `/api/v1/forms` | form slug | | Runbooks | `/api/v1/runbooks` | ID or slug | | Catalog types | `/api/v1/catalog/types` | slug | | Alert sources | `/api/v1/alert_sources` | endpoint path | | Alert routing rules | `/api/v1/routing_rules` | priority within its scope | | Webhooks | `/api/v1/webhooks` | ID | | Service keys | `/api/v1/api_keys` | prefix | | Agents | `/api/v1/agents` | slug | Each takes `GET` and `POST` on the collection, and `PATCH` and `DELETE` on one member. Incident forms are the exception, since the four forms always exist and cannot be created or removed. `GET /api/v1/forms/:slug` reads one, where the slug is `declare`, `update`, `resolve` or `cancel`, and `PATCH` changes one field on it. Creating a severity needs a name. Position is what orders them, first being the most severe, so a new top severity takes position 1. Leave it out to add one at the bottom. Severities, statuses, incident types and incident roles all take `position` the same way, on create and on update, and a position past either end lands on that end. ```bash curl -X POST https://app.firefight.app/api/v1/severities \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "name": "SEV0", "position": 1, "color": "#e5484d", "description": "Everything is on fire." }' ``` **Renaming never moves the slug**, because the slug is what your existing incidents point at. Send `PATCH` with the slug in the path and whatever you want changed. ```bash curl -X PATCH https://app.firefight.app/api/v1/severities/sev0 \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "name": "SEV0 Critical", "enabled": false }' ``` `enabled: false` retires an entry without deleting it. It keeps its slug, stops being offered to responders, and the incidents already using it are untouched. `DELETE` is refused with a `422` while anything still points at the entry, and the message says how many, so disabling is what you want in that case. A listing leaves disabled entries out, matching what responders are offered. Add `?include_disabled=true` when you are managing the list and need to see something in order to turn it back on. A status also needs `lifecycle_stage`, one of `triage`, `active`, `closed` or `canceled`, which is what decides whether that status means the incident is live or over. ### Custom fields, forms and runbooks Custom field options are matched by label, so resending a list renames rather than replaces, and the incidents already holding an option keep pointing at it. See [Custom fields](https://firefight.app/docs/customization/custom-fields.md). A form is changed one field at a time. Name the field with either `custom_field` or `system_field`, then send what you want changed. ```bash curl -X PATCH https://app.firefight.app/api/v1/forms/declare \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "custom_field": "affected_service", "visible": true, "required": true }' ``` Sending `conditions` replaces the conditions on that field rather than adding to them, and passing `[]` clears them. Severity and Status refuse to be hidden or made optional, exactly as they do on the dashboard. See [Incident forms](https://firefight.app/docs/customization/incident-forms.md). Runbook steps and conditions are only touched when you send them, so changing a summary never silently clears the procedure. ### Alert routing rules A routing rule belongs to a scope, which is either the workspace or one alert source, and is addressed by its priority within that scope. Pass `source` with the alert source name to work on one source, and leave it out for the workspace rules that every source without its own falls back to. Writing the first rule for a source is what gives that source its own set of rules. New rules go on the end, and deleting one leaves a gap rather than moving the rules below it up, so read the list back before addressing another rule by priority. `POST /api/v1/routing/evaluate` is a dry run. It takes hypothetical alert fields and returns the rule that matched, the outcome, a per-condition trace, and warnings when a routing role your rules depend on is not set on any catalog attribute. Nothing is created and nobody is notified. See [Testing your routing](https://firefight.app/docs/alerts/testing-routing.md). ## Postmortems | Endpoint | What it does | |---|---| | `GET /api/v1/incidents/:id/postmortem` | The write-up, as HTML, with its status and who wrote it | | `POST /api/v1/incidents/:id/postmortem` | Opens the postmortem, empty or drafted by AI | | `PATCH /api/v1/incidents/:id/postmortem` | Replaces the body, changes the status, or both | Pass `generate: true` when creating to have Firefight draft the first version from the incident. That takes a moment, so the response comes back with a `generation_state` to poll on. Read it back until the state clears. ```bash curl -X POST https://app.firefight.app/api/v1/incidents/3f8c1a90-2b7e-4d15-9c04-6ea5b3d71f28/postmortem \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "generate": true }' ``` Sending `html` replaces the whole document rather than appending, so read the current one first if you mean to add to it. Every version is kept. `status` moves it to `draft`, `in_progress`, `in_review` or `completed`, and needs no version. Reading a postmortem gives you a `version`, and sending a body means sending that version back with it. If somebody edited the postmortem in between, your write is refused with a `409` rather than throwing their work away, and you read again and reapply your edit. A body sent with no version at all is refused with a `422`, since there is no way to tell whether it was built on the current document. ```bash curl -X PATCH https://app.firefight.app/api/v1/incidents/3f8c1a90-2b7e-4d15-9c04-6ea5b3d71f28/postmortem \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "html": "

What happened

The pooler ran out.

", "version": 3 }' ``` A refusal carries no version of its own, on purpose. The version that beat you is not one you can write against, since you have not seen what it says. Starting one is refused with a `422` while the incident is still open, for a canceled incident, and for an incident that already has one. The message says which. All three endpoints check `incidents`. ## Reading an incident's conversation `GET /api/v1/incidents/:id/transcript` returns what people said in the incident channel, oldest last, each message with who said it, when, the text, and whether anything in it was redacted. This checks the **Incident Transcripts** ability rather than `incidents`, and also needs the workspace to have turned transcript access on. Without either it comes back as a `403` naming which is missing. See [Incident conversations](https://firefight.app/docs/workspace/incident-conversations.md). Pass `limit` for up to 500 messages, 100 by default. The response carries a `more_before` cursor, which you pass back as `before` to walk further into the past, and which comes back empty once there is nothing older. ### Agents and service keys `/api/v1/agents` and `/api/v1/api_keys` check admin-only permissions that cannot be granted to anyone, so an admin's personal token reaches them and a service key or agent never can. An agent cannot create another agent or widen its own access. Creating either one returns its token in that response and never again. ```bash curl -X POST https://app.firefight.app/api/v1/agents \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "name": "Support agent", "description": "Triages tickets and opens incidents." }' ``` ```json { "id": "3f8c1a90-2b7e-4d15-9c04-6ea5b3d71f28", "name": "Support agent", "slug": "support_agent", "description": "Triages tickets and opens incidents.", "enabled": true, "granted_abilities": 0, "tokens": [{ "prefix": "ff_4kWm2xPqR", "created_at": "2026-08-26T14:02:00Z", "last_used_at": null }], "token": "ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" } ``` A prefix is the first twelve characters of the token, which is how you name one to revoke it. The new agent holds no abilities at all. Grant it what it needs under [Permissions](https://firefight.app/docs/gateway/permissions.md), or with the `/api/v1/grants` endpoints. `POST /api/v1/agents/:slug/rotate` issues a second token while the first keeps working, so the agent stays up while you update its configuration. `DELETE /api/v1/agents/:slug/tokens/:prefix` ends one when you are ready. See [Agents](https://firefight.app/docs/gateway/agents.md). Sending `permissions` to `/api/v1/api_keys` replaces the whole set rather than adding to it, so read the current one from `GET /api/v1/api_keys` first. A listing never carries a token. ## Worked example: declare, then resolve First fetch the reference data you need. Severities, statuses, and types are specific to your workspace, so their IDs are too. ```bash curl https://app.firefight.app/api/v1/severities \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" ``` ```json { "severities": [ { "id": "c7e2f9a1-3b4d-4c6e-8f0a-5d6b7c8e9f01", "name": "SEV1", "slug": "sev1", "rank": 2, "position": 1, "is_default": false }, { "id": "a1b2c3d4-5e6f-4a7b-8c9d-0e1f2a3b4c5d", "name": "SEV2", "slug": "sev2", "rank": 1, "position": 2, "is_default": true } ] } ``` `position` is the order you set, first being the most severe. `rank` is derived from that order, so a higher rank means a more severe incident. You cannot set rank directly. Now declare the incident. Only `name`, `severity_id`, and `idempotency_key` are required. Omitting `status_id` uses your workspace's default status. You can also pass `incident_type_id`, `declared_by_id` with a member's ID, `visibility` as `public` or `private`, `custom_fields` as a key-value object, and a free-form `source` string that tells responders where the incident came from. Custom fields are checked against your declare form, so an unknown key, an invalid value, or a missing required field returns a `validation_error` instead of creating the incident. ```bash curl -X POST https://app.firefight.app/api/v1/incidents \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "idempotency_key": "datadog-monitor-771204-1626601964", "name": "Checkout latency spike in eu-west", "summary": "p99 latency on the checkout service exceeded 4s.", "severity_id": "a1b2c3d4-5e6f-4a7b-8c9d-0e1f2a3b4c5d", "source": "datadog" }' ``` The response is `201 Created` with the full incident. ```json { "incident": { "id": "0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70", "identifier": "INC-42", "name": "Checkout latency spike in eu-west", "summary": "p99 latency on the checkout service exceeded 4s.", "status": { "id": "b1f6a2d4-8c3e-4f5a-9b7d-1e2c3a4b5d6f", "name": "Investigating", "lifecycle_stage": "active" }, "severity": { "id": "a1b2c3d4-5e6f-4a7b-8c9d-0e1f2a3b4c5d", "name": "SEV2", "rank": 2 }, "type": null, "lead": null, "source": "datadog", "declared_by": null, "declared_at": "2026-07-18T09:12:44Z", "detected_at": null, "resolved_at": null, "created_at": "2026-07-18T09:12:44Z", "updated_at": "2026-07-18T09:12:44Z", "custom_fields": {}, "visibility": "public" } } ``` Update it with `PATCH /api/v1/incidents/:id`. You can change `name`, `summary`, `status_id`, `severity_id`, `incident_type_id`, or assign a lead with `lead_id`, in any combination. Every field in the request lands in one change. Moving the incident to a status in the `closed` stage resolves it, moving it to a `canceled`-stage status cancels it, and moving a resolved or canceled incident back to a `triage` or `active` status reopens it. A resolved incident cannot be canceled directly, and a canceled one cannot be resolved. Reopen it first. The request comes back as `incident_not_active` with the reason. `lead_id` only works while the incident is live, and an incident that has a lead keeps one: sending `lead_id` as `null` is refused. Assigning a lead announces it in the incident channel and tells the person directly, so on a resolved or canceled incident the request comes back as `incident_not_active`. Reopen the incident first if the response is genuinely starting again. ```bash curl -X PATCH https://app.firefight.app/api/v1/incidents/0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70 \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" \ -H "Content-Type: application/json" \ -d '{ "status_id": "f0e1d2c3-b4a5-4968-8776-655443322110", "summary": "Rolled back the 14:02 deploy. Latency recovered." }' ``` The response is `200 OK` with the updated incident, including its `resolved_at` timestamp. :::note Incidents declared through the API behave exactly like incidents declared from Slack. Channels, workflows, and [webhooks](https://firefight.app/docs/api/webhooks.md) all fire the same way. ::: --- Source: https://firefight.app/docs/api/webhooks # Outbound webhooks Webhooks push events out of Firefight as they happen, so your systems react in real time instead of polling the [REST API](https://firefight.app/docs/api/overview.md). Wire them into status pages, ticketing systems, data warehouses, or anything else that should know when an incident moves. ## Creating a subscription Go to **Settings → Webhooks** and add a webhook with a name, a destination URL, and the events it should receive. Managing webhooks requires workspace admin access. The URL must use `http` or `https`, and it must be publicly reachable. Deliveries to addresses that resolve to private networks are refused. Each webhook gets its own signing secret, generated when you create it. Open the webhook's detail panel to reveal the secret and to browse its recent deliveries. Revealing the secret needs admin access. You can also send a test delivery, which replays the newest event from your workspace that the webhook subscribes to, signed exactly like a live delivery. Choose **Send Test** in the webhook's detail panel, or call the `test_webhook` tool over [MCP](https://firefight.app/docs/api/mcp-server.md) with the webhook's id. Either way the delivery appears in the webhook's deliveries list with its response code and timing. A test is refused while nothing the webhook subscribes to has happened in your workspace yet, so declare an incident first on a brand new workspace. ## Events You choose the events per webhook. A webhook only receives what it subscribes to. ### Incident events | Event | Fires when | |---|---| | `incident.created` | An incident is declared | | `incident.updated` | An incident's core fields change | | `incident.accepted` | Someone accepts a triaged incident | | `incident.resolved` | An incident is resolved | | `incident.reopened` | A closed incident is reopened | | `incident.canceled` | An incident is cancelled | | `incident.escalated` | An incident is escalated | | `incident.marked_duplicate` | An incident is marked as a duplicate | | `incident.merged_into` | An incident is merged into another | | `lead.assigned` | An incident lead is assigned | | `role.assigned` | Someone is given an incident role other than lead | | `role.unassigned` | An incident role other than lead is cleared | ### Action events | Event | Fires when | |---|---| | `action.created` | An action item is created | | `action.picked_up` | Someone picks up an action item | | `action.reassigned` | An action item is handed to someone else | | `action.completed` | An action item is completed | ### Runbook events | Event | Fires when | |---|---| | `runbook.attached` | A runbook attaches to an incident | ### Timeline note events | Event | Fires when | |---|---| | `milestone.noted` | Firefight adds a note to an incident's timeline | The payload carries the note's kind, its statement, who said it, when, the message it was read from, and the link to that message. See [Timeline notes](https://firefight.app/docs/incidents/concepts#timeline-notes). ### Postmortem events | Event | Fires when | |---|---| | `postmortem.generated` | A postmortem is generated for an incident | | `postmortem.edited` | A postmortem's content is edited | ### Relationship events | Event | Fires when | |---|---| | `relationship.created` | An incident is linked to a related incident | ## Delivery Firefight sends each event as an HTTP `POST` with a JSON body and the User-Agent `Firefight-Webhooks/1.0`. Your endpoint has 7 seconds to respond. Any `2xx` response counts as success, and anything else, including timeouts and connection errors, counts as failure. Every request carries these headers. | Header | Contents | |---|---| | `X-Webhook-Version` | Payload schema version, currently `1` | | `X-Webhook-Event` | The event type, for example `incident.resolved` | | `X-Webhook-Delivery` | Unique ID of this delivery | | `X-Webhook-Attempt` | Attempt counter for this delivery | | `X-Webhook-Timestamp` | ISO 8601 time the request was signed | | `X-Webhook-Signature` | Signature in the form `v1=` | Failed deliveries are not retried automatically. Each webhook keeps a delivery history in **Settings → Webhooks** showing the event, state, and response code of its recent deliveries, and deliveries are kept for 7 days. Replay any of them from that panel, which needs admin access. A replay sends the exact bytes that were sent the first time, so it is the same delivery arriving again rather than a fresh render of an incident that may have changed since. That means the `delivery_id` in the body is the original one, and a consumer that has already processed it can recognise the repeat and ignore it. If a webhook fails 10 times in a row over the span of an hour or more, Firefight deactivates it so a dead endpoint does not accumulate failures forever, and posts a notice to your incidents channel naming the webhook. Fix the endpoint, then re-enable the webhook in **Settings → Webhooks**. Any successful delivery resets the failure counter. ## Payload Every payload shares the same envelope, with event-specific detail under `data`. This is an `incident.created` delivery. ```json { "id": "3c8e5f2a-9b1d-4e7c-a6f0-8d2b4c6e9a13", "version": "1", "event_type": "incident.created", "workspace_id": "7a1c3e5b-2d4f-4a68-9c0e-1b3d5f7a9c2e", "occurred_at": "2026-07-18T09:12:44Z", "data": { "incident": { "id": "0d9b2c1e-7f3a-4b8e-9c5d-2a6e8f1b4d70", "identifier": "INC-42", "name": "Checkout latency spike in eu-west", "summary": "p99 latency on the checkout service exceeded 4s.", "status": { "id": "b1f6a2d4-8c3e-4f5a-9b7d-1e2c3a4b5d6f", "name": "Investigating", "lifecycle_stage": "active" }, "severity": { "id": "a1b2c3d4-5e6f-4a7b-8c9d-0e1f2a3b4c5d", "name": "SEV2", "rank": 2 }, "type": null, "lead": null, "source": "datadog", "declared_by": { "id": "e4a8b2c6-1d3f-4e5a-b7c9-0f1a2b3c4d5e", "type": "user", "name": "Maja Kovac", "email": "maja@example.com" }, "declared_at": "2026-07-18T09:12:44Z", "detected_at": null, "resolved_at": null, "created_at": "2026-07-18T09:12:44Z", "updated_at": "2026-07-18T09:12:44Z", "custom_fields": {}, "channel_id": "C0912345678", "channel_name": "inc-42-checkout-latency-spike" }, "actor": { "id": "e4a8b2c6-1d3f-4e5a-b7c9-0f1a2b3c4d5e", "type": "user", "name": "Maja Kovac", "email": "maja@example.com" } } } ``` The `actor` is whoever caused the event. Its `type` is `user` for a person and `api_key` for an integration, in which case `email` is `null`. Action, postmortem, and relationship events carry their own detail under `data` alongside the incident. ## Verifying signatures Every delivery is signed with the webhook's secret using HMAC-SHA256. The signed input is the scheme, the timestamp, and the raw body joined with colons. ``` v1:: ``` To verify, compute the HMAC-SHA256 hex digest of that string with your signing secret and compare it against the digest in `X-Webhook-Signature` using a constant-time comparison. Reject the delivery if it does not match. ```js const crypto = require("crypto") function verify(secret, timestamp, rawBody, signatureHeader) { const expected = "v1=" + crypto .createHmac("sha256", secret) .update(`v1:${timestamp}:${rawBody}`) .digest("hex") return crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(signatureHeader)) } ``` :::note Check that `X-Webhook-Timestamp` is recent, for example within 5 minutes, before comparing signatures. That blocks replay of captured requests. Legitimate replays from the delivery history are re-signed with a fresh timestamp. ::: :::tip Use an `https` URL and treat the signing secret like a password. Each webhook has its own secret, so a leak only affects that one subscription. Delete and recreate the webhook to rotate it. ::: --- Source: https://firefight.app/docs/catalog/entries-and-attributes # Entries, attributes & relationships Everything in the catalog follows one shape. A type defines a schema of attributes, and entries fill that schema in. This page walks through defining both, plus the relationships that connect entries and the API path for keeping entries in sync from outside tools. ## Creating and editing types On the **Catalogue** page, **Create Type** opens a dialog where you name the type, describe it, pick an icon and a color, and define its attributes. On a type's own page, **Edit Type** opens the same dialog for changes. You can edit built-in types too, including adding attributes to them. ## Attributes Each attribute has a name, a kind, and an optional required flag. These are the kinds you can choose. | Kind | What it holds | |---|---| | Text | Free-form text, such as a description or a repository URL | | Number | A numeric value | | Boolean | A yes or no flag, such as Is Production | | Select | One choice from a fixed list of options you define | | List | A list of values | | Reference | A link to an entry of another catalog type you pick, such as a service's Owner Team | | Slack Channel | A Slack channel | | Member | One workspace member | | Members | Multiple workspace members | A Select attribute requires at least one option. A Reference attribute requires you to pick which type it points at. An attribute's kind is fixed after creation, so if you need a different kind, add a new attribute instead. ## Routing roles Some attributes do a job for [alert routing](https://firefight.app/docs/alerts/routing-rules.md): when a rule pages a team, Firefight reads the people from the attribute marked as **Members**, the accountable person from **Manager**, and the channel to post in from **Notification channel**. Member and Slack Channel attributes show a **Role** dropdown in the type editor where an admin picks that job. Each role can be held by one attribute per type. The built-in Team and Service types come with their roles already set, so routing works out of the box. The role follows the attribute, so renaming an attribute never breaks routing. When a role a rule depends on is not set on any attribute, the Alert Routing page shows a warning naming the gap. The built-in types ship with these attributes. | Type | Default attributes | |---|---| | Team | Description, Slack Channel, Manager, Members | | Service | Description, Owner Team (references Team), Tier (Critical, Standard, or Internal), Repository, Slack Channel | | Environment | Description, Is Production, Region | | Functionality | Description, Owner Team (references Team) | ## Entries On a type's page, **Add** opens a form with a field for the entry's name and one field per attribute. Entries appear in the table with their attribute values as columns. Clicking an entry opens a detail panel showing every attribute, when the entry was created and last updated, and actions to edit or delete it. Each entry gets a slug derived from its name, lowercased with underscores, such as `payment_service` for "Payment Service". The slug is the entry's stable identifier. It cannot change after creation, it must be unique within the type, and it is what alerts and the API match against. :::note Alert enrichment matches an alert's `service` field against service entry slugs. Name your entries so their slugs line up with what your monitoring tools send, or configure your tools to send the slug. ::: ## Relationships Relationships between entries come from Reference attributes. When you set a service's Owner Team, Firefight stores a relationship from that service to that team. Clearing the attribute removes the relationship. These links are what make the catalog a map rather than a list. The owning-team features in alert routing follow the service to team relationship, and connected AI agents see each entry's relationships when they search the catalog. See [Putting the catalog to work](https://firefight.app/docs/catalog/using-catalog.md). ## Entries managed by external tools If a service registry, an infrastructure repo, or a script is the source of truth for part of your catalog, push entries through the API instead of maintaining them by hand. When a create request includes a `source` (the name of the pushing system) and an `external_id` (its identifier there), Firefight treats the pair as the entry's identity. Pushing the same pair again updates the existing entry instead of creating a duplicate, so a sync job can run repeatedly and stay idempotent. Both values must be provided together. API keys are minted under **Settings → API Keys** and need catalog scopes for these endpoints. ## Reading and writing the catalog | Request | What it does | |---|---| | `GET /api/v1/catalog/types` | List types with their attribute definitions | | `POST /api/v1/catalog/types` | Create a type with its attributes | | `GET /api/v1/catalog/types/:slug` | Fetch one type | | `PATCH /api/v1/catalog/types/:slug` | Change a type or its attributes | | `DELETE /api/v1/catalog/types/:slug` | Delete a type and its entries | | `GET /api/v1/catalog/types/:slug/entries` | List a type's entries | | `POST /api/v1/catalog/types/:slug/entries` | Create an entry, or update it when `source` and `external_id` match an existing one | | `GET /api/v1/catalog/entries/:id` | Fetch one entry | | `PATCH /api/v1/catalog/entries/:id` | Update an entry's name or attributes | | `DELETE /api/v1/catalog/entries/:id` | Delete an entry | Attributes are matched by name when you write a type, so resending a list renames rather than replaces and the entries already holding a value keep it. Leaving one out removes it, which is refused while an active entry still uses it. Send no attributes at all and the shape is left alone, so changing a description is safe. A reference attribute names the type it points at by slug, and reading a type back gives you that slug, so a type can be read and written without looking up an ID. Built-in types keep their slug and their own attributes. Deleting a type is refused while another type's reference attribute points at it, since its entries would vanish from that picker, and the refusal names the attribute in the way. --- Source: https://firefight.app/docs/catalog/overview # Service catalog The catalog is Firefight's answer to the two questions every incident starts with. What is this thing, and who owns it? You record your services, teams, and environments once, and Firefight uses that map everywhere. Alerts route to the owning team, incident forms offer your real services as options, and connected AI agents can answer "who owns checkout?" without anyone digging through wikis. The catalog lives at **Catalogue** in the dashboard sidebar. The main page shows a grid of your catalog types with an entry count on each card. Click a type to see its entries in a table, add new ones, or open an entry to inspect its attributes. ## The four built-in types Every workspace starts with four types. They are the vocabulary the rest of Firefight understands, so alert routing and catalog enrichment key off them directly. | Type | What it holds | |---|---| | Team | Teams that own and operate services, with a Slack channel, a manager, and members | | Service | Services and applications in your infrastructure, each with an owner team, a tier, and a repository | | Environment | Deployment environments, with a production flag and a region | | Functionality | Business capabilities and product features, each with an owner team | Built-in types come with a sensible set of attributes out of the box, and you can add more. Their identity is fixed so that routing and enrichment always know where to look, but everything else about them is yours to shape. See [Entries, attributes & relationships](https://firefight.app/docs/catalog/entries-and-attributes.md) for the full attribute list on each type. ## Custom types The four built-ins rarely cover everything. Click **Create Type** on the Catalogue page to add your own, such as Vendor, Database, or Feature Flag, or create one over the [API](https://firefight.app/docs/catalog/entries-and-attributes#reading-and-writing-the-catalog) or by asking an agent connected over [MCP](https://firefight.app/docs/api/mcp-server.md). A custom type gets a name, a description, an icon, a color, and any set of attributes you define, including references to entries of other types. Custom types behave exactly like built-in ones in the dashboard and the API. :::tip Start small. A catalog with ten real services and their owning teams already powers owning-team alert routing. You can grow it from there. ::: ## What the catalog unlocks A filled-in catalog is not documentation for its own sake. It drives concrete behavior. - Alert routing can invite the owning team of the affected service and notify its channel. See [Putting the catalog to work](https://firefight.app/docs/catalog/using-catalog.md). - Custom fields on incidents can offer catalog entries as their options, so "Affected service" is always a real service. - The API and connected AI agents can read the catalog, so external tools and agents share the same map you do. --- Source: https://firefight.app/docs/catalog/using-catalog # Putting the catalog to work The catalog pays off in three places. Alerts route by ownership, incident fields stay structured, and the same map is readable by the API and by AI agents you connect. This page covers each payoff and what the catalog needs to contain for it to work. ## Alert routing by ownership When an alert arrives carrying a `service`, `team`, `environment`, or `functionality` field that names a catalog entry by slug, Firefight enriches the alert with catalog context before routing rules run. - Related entries one hop away are merged in. An alert with only `service: checkout` also gets `team` filled from the service's owner team. - The entry's attributes become dotted fields, so a rule condition like `service.tier` `is one of` `Critical` routes by catalog metadata the alert never sent. - When you write a condition on a catalog field, the value picker offers your catalog entries directly. Routing outcomes can then target ownership instead of hardcoded names. | Outcome | What the catalog resolves | |---|---| | Notify the owning team's channel | The service's own Slack Channel attribute wins when set, otherwise the owner team's channel | | Invite the owning team | The owner team's members and manager are invited to the incident channel | Resolution happens at fire time, not when you save the rule. Reorganize ownership in the catalog and every rule follows without edits. When something cannot be resolved, for example a service with no owner team or a team with no channel set, the incident is still created and the gap is recorded as a note on it. See [Routing rules](https://firefight.app/docs/alerts/routing-rules.md) for the full rule reference. :::tip For owning-team routing you need exactly three things. The alert's `service` field matches a service entry's slug, that service has its Owner Team set, and the team has its Slack Channel, Members, or Manager attributes filled in. ::: ## Catalog-backed custom fields Custom fields under **Settings → Custom Fields** can draw their options from the catalog instead of a hand-maintained list. - A **Catalog reference** field stores one catalog entry, and a **Catalog multi-reference** field stores several. Point the field at a type, such as Service, and responders pick from your real services. - Single-select and multi-select fields can also choose **From catalogue** as their option source, using a type's entries as the choices. Attach these fields to the incident lifecycle forms under **Settings → Forms** and an "Affected services" field becomes structured data that always matches the catalog, instead of free text that drifts. ## Readable by agents and the API The catalog is part of Firefight's readable surface for external tools. - The REST API exposes types and entries under `/api/v1/catalog/`, listed in [Entries, attributes & relationships](https://firefight.app/docs/catalog/entries-and-attributes.md). - AI agents connected over MCP get a catalog search tool. An agent investigating an incident can ask which team owns a service, find its Slack channel, and follow relationships between entries, with member attributes resolved to display names. Connections are managed under **Settings → API Keys**. An agent can also shape the catalog, creating types and writing entries, but only if you granted it that. Read and write are separate abilities, so an agent that only looks things up cannot change them. Because agents, the API, alert routing, and incident fields all read the same entries, keeping the catalog current in one place keeps every consumer current at once. --- Source: https://firefight.app/docs/customization/custom-fields # Custom fields Custom fields let you capture the data your team cares about on every incident, things like affected services, impacted environment, or customer segment. You define a field once in **Settings → Custom Fields**, then attach it to one or more of the [incident forms](https://firefight.app/docs/customization/incident-forms.md) so responders fill it in at the right moment. ## Field types Click **Add field** to create a field. Give it a name, an optional description that responders see as help text, and pick a type: | Type | Best for | | --- | --- | | Text | Short or long-form text input | | Number | Numeric input for counts or estimates | | Link | An external reference URL | | Single-select | One choice from a curated set | | Multi-select | Multiple choices from a curated set | | Catalog reference | Selecting one catalog entry | | Catalog multi-reference | Selecting multiple catalog entries | Each field gets a stable key derived from its name. The key is shown next to the field in the list and does not change afterward, so integrations and reports can rely on it even if you rename the field. ## Option sources Select and reference fields need a source for their options: | Option source | How it works | | --- | --- | | Fixed list | You add the options yourself, one input per option | | From catalogue | Options come from a catalog type you pick, such as Services or Teams | Text, number, and link fields take free input and have no options. Single-select and multi-select fields can use either a fixed list or a catalog type. Catalog reference and catalog multi-reference fields always draw from a catalog type. :::tip Prefer catalog-backed options over fixed lists when the values represent real things like services or teams. The options stay in sync with your catalog automatically, the answers stay structured, which makes them far easier to report on later, and a long list stays manageable in a way a hand-typed one does not. ::: ### Choose the type before the field is used A field's type and option source can be changed freely up until the first incident stores a value in it. After that both lock, because changing them would reinterpret every answer already recorded. The two dropdowns grey out and hovering either one tells you how many incidents are involved. Changing the type before that point clears the options, since a list written for one shape rarely fits another. If a field in use turns out to be the wrong shape, turn it off and add a new one. The old field keeps its answers on the incidents that have them, and you attach the replacement to your forms. So it is worth a moment at creation deciding whether a value is really a number, a single choice, or several. ## Where values get filled in A custom field only appears to responders once you attach it to a form. There are four, one for each moment in an incident's life: | Form | Opens with | | --- | --- | | Declare | `/ff new` | | Update | `/ff update` | | Resolve | `/ff resolve` | | Cancel | `/ff cancel` | Each one also opens from the dashboard, so a field you attach is asked there too. See [Incident forms](https://firefight.app/docs/customization/incident-forms.md) for how attaching, ordering, and required settings work. Values captured this way appear on the incident page in the web dashboard in a Custom Fields panel, and they are included in the incident record that [AI postmortem generation](https://firefight.app/docs/postmortems/generating-with-ai.md) draws from. ## Working with a fixed list Each option in a fixed list gets its own input. Click **Add option** to add one, drag the handle to change the order responders see in the dropdown, and use the switch to turn an option off. Renaming an option is safe at any time. Incidents that already hold it, and any runbook conditions that match on it, follow the new name. Nothing needs updating afterward. Turning an option off removes it from the dropdown without touching the incidents that already hold it. Those keep showing the option and reading normally, responders just cannot pick it on new incidents. Turn it back on whenever you want it available again. Deleting is only possible while nothing points at the option. Once an incident holds it or a runbook condition matches on it, Delete is disabled and hovering it tells you how many references there are. That count covers both, so an option no incident has ever used can still be held by a condition. Turn the option off instead, which is what you want in almost every case where a value has gone out of use. ## Turning a field off, and deleting one A field follows the same rule as its options, one level up. **Turning a field off** stops it being collected without disturbing anything already recorded. It disappears from every form it is attached to, and the incidents holding values keep them. The switch in the **Enabled** column does this, and turning it back on puts the field back on the same forms. **Deleting a field** is only possible while nothing depends on it. Delete is disabled while any form is attached to the field, and disabled again while any incident holds a value for it, because those answers are part of that incident's history. Hovering tells you which of the two is blocking, and how many. So a field that has been used in anger gets turned off rather than deleted. That is not a limitation to work around, it is what keeps a resolved incident readable a year later. ## Managing fields The list in **Settings → Custom Fields** shows every field with its key, its type, its catalog type if it has one, how many forms it is attached to, and whether it is enabled. The menu at the end of each row edits the field or deletes it. ## Over the API and MCP Fields can be created, changed and removed at `/api/v1/custom_fields`, and with the `upsert_custom_field` MCP tool. Options are matched by label, so resending a list renames rather than replaces, and the incidents already holding an option keep pointing at it. See [Using the API](https://firefight.app/docs/api/using-the-api#custom-fields-forms-and-runbooks). --- Source: https://firefight.app/docs/customization/incident-forms # Incident forms Incident forms are the dialogs responders fill in at key moments of an incident's life. Firefight ships four of them with sensible defaults, and **Settings → Forms** lets you tune each one, which fields appear, in what order, and which are required, without any responder retraining. Slash commands simply start showing the updated dialog. ## The four lifecycle forms | Form | When responders see it | | --- | --- | | Declare | When someone declares an incident with `/ff new` | | Update | When someone posts a status update with `/ff update` | | Resolve | When the incident is resolved or closed with `/ff resolve` or `/ff close` | | Cancel | When an incident is cancelled with `/ff cancel` as not a real incident | The same four forms open from the dashboard, so whatever you configure here is asked on both surfaces. See [Running an incident from the dashboard](https://firefight.app/docs/incidents/running-from-the-dashboard.md). Every form works out of the box with built-in fields such as the incident's name, severity, and summary. You only need to touch **Settings → Forms** when you want to change what gets asked. The Cancel form is the exception. Nothing on it is switched on to begin with, so `/ff cancel` dismisses a false positive in one step with no dialog at all. Turn something on, or add a custom field such as "Cancellation reason", and the dialog starts appearing. ## Editing a form Pick a form from the list on the left. The editor shows each field the way responders will see it, with a preview of its input. For every field you can: - **Reorder** by dragging the handle next to the field. - **Show or hide** with the Visible toggle. Hidden fields stay configured but disappear from the dialog. - **Require or make optional** with the Required toggle. A few essential fields are locked as required and cannot be hidden or made optional. Some fields are listed switched off rather than missing, so you can see everything a form is able to ask before deciding to ask it. A field can also carry a line explaining that responders will not be asked it yet, which happens when there is genuinely nothing to choose. ## Conditional fields Fields can appear only when they are relevant. Add a condition to a field and it only shows when the incident's type or severity matches. Conditions support "is one of" and "is not one of" against the incident types and severities you have configured. :::tip Use conditions to keep the Declare form short. A "Customer communication owner" field that only appears for your highest severities keeps low-stakes declarations fast while still capturing what matters when it counts. ::: ## Adding custom fields Click **Add custom field** on a form to attach any field defined in **Settings → Custom Fields**. The same field can be attached to several forms, so a value requested at declare time can be revisited at resolve time. Attached custom fields can be reordered, hidden, required, conditioned, and removed just like you would expect. Removing a custom field from a form does not delete the field definition or any values already captured. If the field you need does not exist yet, create it first. See [Custom fields](https://firefight.app/docs/customization/custom-fields.md) for field types and option sources. ## Over the API and MCP `GET /api/v1/forms/:slug` reads one form and every field on it, hidden ones included, and `PATCH` changes one field at a time. The `get_form` and `upsert_form_field` MCP tools do the same. Sending conditions replaces the set on that field rather than adding to it, and the four forms cannot be created or removed. See [Using the API](https://firefight.app/docs/api/using-the-api#custom-fields-forms-and-runbooks). --- Source: https://firefight.app/docs/gateway/activity # Activity Activity is the record of everything that passed through the gateway. If an agent touched one of your systems, a person changed how the workspace is set up, or a request was refused, it is here. Go to **Gateway → Activity**. The page requires admin access. ## What is recorded | Recorded | Not recorded | |---|---| | Every call to a connected tool, reads included, by anyone | Reading incidents, alerts and the catalog inside Firefight | | Every change to workspace configuration | Declaring, updating and closing incidents | | Every refused request, and every request waiting for approval | | Calls that leave Firefight are always recorded, even when they only read, because "what did the agent look at" is the question this page exists to answer. Reading Firefight's own data is not, since it happens constantly and tells you nothing. Incident participation by a person is recorded on the incident itself, on its timeline, where it belongs next to what happened. See [Incident concepts](https://firefight.app/docs/incidents/concepts.md). ## Reading the log Each row is one request, newest first, showing the most recent 200. | Column | What it tells you | |---|---| | When | When the request was made | | Principal | Who made it, a person, an API key, or an agent | | Source | Where it came from. Dashboard, Slack, API, or MCP | | Action | What was asked for, as the capability's name | | Decision | `allow`, `deny`, or `pending` | | Outcome | Whether an allowed action succeeded or failed, with the error when it failed | | Duration | How long the action took to run | **Decision** is the gateway's answer. `deny` means the principal did not hold the permission and nothing ran. `pending` means the request is waiting for approval, and a second row appears when it is decided and runs. See [Approvals](https://firefight.app/docs/gateway/approvals.md). **Outcome** is what happened after an `allow`. A row that was allowed but shows no outcome yet is still running. Use the filter at the top right to show only one decision. Filtering to `deny` is the quickest way to see what an agent tried and could not do, and filtering to `pending` shows what is waiting. ## What to do with it Look here when an agent did something unexpected, when you are checking what a new integration has been used for, or when someone asks who changed a routing rule. The principal and source columns answer who and how, and the action and outcome columns answer what. An agent's own calls appear here like everyone else's, under the name of the key it connected with. The same log is available at `/activity` over the API and through the `search_activity` MCP tool, to an admin's personal token. See [Connect AI agents](https://firefight.app/docs/api/mcp-server.md). --- Source: https://firefight.app/docs/gateway/agents # Agents An agent is an AI that takes part in your workspace under its own name. It declares incidents, raises and picks up work, pulls people in and reports what it found, and every one of those actions is recorded as the agent's, not as the work of whoever set it up. An agent holds only the abilities you grant it, so it starts able to do nothing. Go to **Gateway → Agents**. The page requires admin access. Everything on this page is also available over the [API](https://firefight.app/docs/api/using-the-api#agents-and-service-keys) and as [MCP tools](https://firefight.app/docs/api/mcp-server#agents-and-keys), to an admin. Both check a permission that cannot be granted to anyone, so an agent can never create another agent or issue itself a new token. ## Creating an agent Give it a name people will recognise on a timeline, such as Support agent, and a short line saying what it does. The slug is the short machine-readable name, filled in from what you type, and it is fixed once the agent exists because grants and the activity log record it. Creating the agent hands you its token once, on the screen. Copy it then. It is never shown again, and if you lose it you issue a new one rather than recovering the old. Configure your agent to send that token as a Bearer header, the same way any other Firefight credential works. ``` Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA ``` A brand new agent can authenticate and nothing else. The row says **None granted, so it can do nothing** until you give it something under [Permissions](https://firefight.app/docs/gateway/permissions.md), where it appears alongside your people and service keys. ## What to grant it Start with what the agent actually needs and add to it. An agent that triages support tickets and opens incidents needs `incidents: read` and `incidents: create`. One that also works the incident needs `incidents: update`, which covers raising work, taking it, finishing it, escalating, inviting and linking. Reaching a connected tool is always a separate grant, and can be limited to particular environments. See [Permissions](https://firefight.app/docs/gateway/permissions.md) for how grants and environments work, and [Approvals](https://firefight.app/docs/gateway/approvals.md) for making a risky ability wait for a person. An agent never inherits anyone's access. Creating it as an admin does not make it an admin. ## Tokens An agent can hold more than one token at a time, which is what makes rotation safe. **Issue a new token** hands you a second working credential while the first keeps working, so you update the agent's configuration on your own schedule and nothing goes down in the middle. When the agent is running on the new one, open **Tokens** on its row and revoke the old. Revoking a token stops it immediately. Nothing else moves: the agent keeps its name, its abilities and everything it has already done, because those belong to the agent rather than to a secret. The Tokens dialog lists every working token with when it was issued and when it was last used, so you can tell which one is live before revoking anything. An agent with no working token cannot act, and its row says **No token** rather than looking like it is running. ## Disabling and deleting Disabling an agent stops it acting while keeping its row, its slug and its grants. The row stays on the list so you can switch it back on. Deleting an agent takes its tokens and its abilities with it. What it did stays on the timelines it touched, so a past incident still names the agent that worked it. ## Agents and service keys Both are credentials that act without a person behind them, and both start with no access. The difference is what the record says afterwards. | | Agent | Service key | |---|---|---| | Where you make it | **Gateway → Agents**, the API, or MCP | **Settings → API Keys**, the API, or MCP | | Who the timeline names | The agent, by name | The key, by name | | Rotating the credential | Issue a second token, revoke the first | Create a second key, delete the first | | Best for | An AI taking part in incidents | A pipeline or script calling Firefight | Use an agent when something will appear on an incident timeline as a participant. Use a service key when something is plumbing, such as an alerting pipeline that declares incidents and never speaks again. ## Seeing what an agent has done Every governed action an agent takes lands in [Activity](https://firefight.app/docs/gateway/activity.md), including the ones that were refused. Its work inside an incident also shows on the incident's own timeline, where an agent is marked as a machine rather than shown with a person's initials. See [Connect AI agents](https://firefight.app/docs/api/mcp-server.md) for what an agent can actually do once it is connected. --- Source: https://firefight.app/docs/gateway/approvals # Approvals An approval is a request that is allowed in principle but held until someone with the right role says yes. The action does not fail and does not need to be repeated, it waits, and runs by itself once approved. Go to **Gateway → Approvals**. Everyone in the workspace can see the page, and the sidebar shows how many requests are waiting. ## When an action waits for approval Nothing waits for approval until an admin says it should. Out of the box every granted ability runs as soon as it is called. An approval rule changes that for the calls it matches: Firefight parks the call and asks for a decision instead of running it. Rules live under **Gateway → Permissions**, in the **Approval rules** block, and need admin access. They can also be written over the API (`/approval_rules`) and with the `upsert_approval_rule` MCP tool, from an admin's personal token. Each rule answers three questions about the call and three about the decision. | Question | Choices | |---|---| | Which abilities | Pick specific abilities, or leave blank for any | | Which risk levels | Read, write, destructive, or any | | Which environments | Pick environments, or leave blank for any | | Who approves | Anyone with the owner role, anyone with the admin or owner role, or specific people and agents | | Where to ask | The incidents channel, a direct message to each approver, or both | | Requester can approve their own request | On by default, switch it off to require someone else | | Agents may decide this rule | Appears once an agent or service key is named. Off by default, so the named machines are listed but only the people decide | A call is held by the first rule that matches it, reading top to bottom, so put the narrow rules above the broad ones and use the arrows to reorder. Switch a rule off to keep it without enforcing it. An ability covered by a rule shows an **approval** marker wherever it is handed out, so whoever is granting it knows it will wait. Two things never wait, whatever the rules say: deciding an approval, and changing permissions or the rules themselves. Otherwise a rule could hold the screen you would use to remove it. :::tip Start with one rule for destructive abilities in production, approved by any admin, asked in the channel. That catches the calls that cannot be undone without slowing down the rest. ::: ## What the requester sees Firefight tells the person or agent that made the request that it is waiting, and tells them again once it is decided. Nothing else is needed from them. **In Slack**, the reply says the action needs approval, that the request has been sent, and that they will hear back either way. When it is approved, Firefight carries the action out and posts a message saying who approved it and that it went ahead. If it is denied, the message says who declined and that nothing changed. **On the dashboard**, the page they were on shows the same message. The decision arrives as a direct message in Slack, and the change they asked for is applied without them having to redo it. **Over the API**, the response comes back as accepted rather than completed, with the message `approval_required` and an approval id. Once approved, the caller repeats the identical request with the header `X-Approval-Id` set to that id, and it runs. See [Using the API](https://firefight.app/docs/api/using-the-api.md). **Over MCP**, the tool result says approval is required and gives the approval id. The agent retries the same call with `approval_id` once approved. See [Connect AI agents](https://firefight.app/docs/api/mcp-server.md). An agent connected as you can also be the one watching. With `search_approvals` it sees a request that names you, tells you, and approves it with `approve_approval` once you say so. Connected as you, it decides as you. An agent can also decide under its own name. Give it a service key under **Developer → API Keys**, name that key as an approver on the rule, and switch on **Agents may decide this rule**. From then on the agent, connected with that key, can approve or deny what the rule holds, and every decision is recorded under the agent's name. Nothing asks the agent, it polls `search_approvals` or `GET /approvals?status=pending` and sees the requests that name it. A rule without that switch refuses a machine's click even when the machine is named, so the default stays a person saying yes. An approval covers exactly the request that was made, with the same details. It admits that request once. Changing anything about it, or calling again after it has run, is a new request and asks again. ## Deciding Every pending request is listed under **Gateway → Approvals**, and posted in Slack wherever the rule said to ask: the incidents channel, a direct message to each approver, or both. Every copy carries **Approve** and **Deny** buttons. Decide from any of them, the result is the same. The Slack message names who is asking, what they want to run, the scope and details of the call, and who may decide, either the role required or the people the rule named. Once decided, every copy of the message updates in place to show the verdict and who gave it, so nobody approves the same request twice. On the dashboard, the **Pending approvals** table shows the same information with where the request came from, and Approve and Deny buttons for anyone who may decide. Below it, **Recently resolved** lists the last decisions with their status and who made them. If you press a button on a request you cannot decide, Firefight tells you why. Either the request is no longer pending, you are not one of its approvers, or the rule asks for someone other than the requester. ## Approving your own request By default you can approve a request you raised yourself. That is deliberate. When your AI agent proposes an action and you approve it, you are the safety check, and asking a second person to confirm what you already read would only slow the incident down. Switch **Requester can approve their own request** off on a rule and approval must come from someone else. The dashboard and Slack both refuse the requester's own click in that case, with a message saying so. ## Permission to decide Approving or denying needs the `approvals` permission, held by admins and owners. It can be granted to a member like any other permission, and to an agent or service key, which then also needs to be named on the rule with **Agents may decide this rule** switched on. See [Permissions](https://firefight.app/docs/gateway/permissions.md). Every request and its outcome is also recorded in [Activity](https://firefight.app/docs/gateway/activity.md), including the ones that are still waiting. --- Source: https://firefight.app/docs/gateway/overview # The gateway The gateway is the part of Firefight that stands between anyone acting, a person, an API key or an AI agent, and anything that action touches. It answers three questions, and the dashboard has one page for each. | Question | Page | Who sees it | |---|---|---| | Who may do what, and where | [Permissions](https://firefight.app/docs/gateway/permissions.md) | Admins and owners | | What is waiting for someone to say yes | [Approvals](https://firefight.app/docs/gateway/approvals.md) | Everyone | | What was done, by whom, and whether it was allowed | [Activity](https://firefight.app/docs/gateway/activity.md) | Admins and owners | All three live under **Gateway** in the dashboard sidebar, and everything on them is also reachable over the [API](https://firefight.app/docs/api/overview.md) and [MCP](https://firefight.app/docs/api/mcp-server.md), so an agent can administer access the same way an admin does. Approvals is first, and carries a count of requests waiting on a decision, so a responder sees at a glance whether anything needs them. ## Why one gate An AI agent connected to your workspace can read a database, open a ticket or change a routing rule through the same connections your team uses. Giving it a login to each tool means trusting it with everything that login can do. The gateway is how you avoid that. You connect the tool once, enable the capabilities you want available, and then grant them. A grant says who may call a capability and in which environments. An agent holds only what you gave it, never what its creator can do. When a capability is sensitive, an approval rule can hold the call until a person approves it. And whatever happens, it is written down. The same gate governs Firefight's own settings. A member who is not an admin can be given runbooks or webhooks without being given everything else, and the change they make is recorded like any other. ## What the gateway covers The gateway checks every action a person or machine takes outside of responding to an incident. - **Connected tools.** Every capability from an integration, whether called by a person on the dashboard or by an agent over MCP. - **Workspace configuration.** Statuses, severities, types, incident roles, custom fields, forms, runbooks, the catalog, alerts and alert sources, routing rules, and webhooks. - **Approvals themselves.** Approving or denying a request is an action with its own permission. Declaring, updating and closing incidents is something every member can do. The gateway does not stand in the way of that, and those changes are recorded on the incident's timeline rather than in Activity. --- Source: https://firefight.app/docs/gateway/permissions # Permissions Permissions decide who can use the capabilities your [integrations](https://firefight.app/docs/api/integrations.md) provide, and where. Everything granted here is enforced by the gateway, whether the request comes from Slack, the dashboard, the API or an AI agent, and every governed action lands in [Activity](https://firefight.app/docs/gateway/activity.md). Enabling a capability makes it available, granting it decides who may actually call it. The same grants cover Firefight's own settings, so a member can be given runbooks or webhooks without being made an admin. Go to **Gateway → Permissions**. The page requires admin access. Everything on it is also available over the [API](https://firefight.app/docs/api/overview.md) (`/abilities`, `/principals`, `/permission_sets`, `/grants`, `/approval_rules`) and as [MCP tools](https://firefight.app/docs/api/mcp-server.md), to an admin's personal token. The resources you grant here line up with what the API and MCP expose, so a permission means the same thing wherever the request came from. ## What everyone starts with Some access needs no grant at all, so check here before granting anything. | Who | What they can do with no grants | |---|---| | Admins and owners | Everything, including every capability your integrations provide | | Members | See everything in Firefight, and declare, update and close incidents | | Agents | Nothing | | Service keys | Nothing | Members can read what is in Firefight and take part in incidents, and that is the same whether they are clicking a button in Slack, using the dashboard, calling the API with a personal token, or working through an AI agent connected as them. Changing how the workspace is set up needs a grant or admin access, and anything that reaches a connected tool needs an explicit grant. Agents and service keys start with nothing at all, so everything they do is something you deliberately gave them, and neither inherits the access of whoever created it. An [agent](https://firefight.app/docs/gateway/agents.md) is an AI acting under its own name, which is what an incident timeline records when it takes part. A service key is for automation with nothing to say on a timeline, such as an alerting pipeline. Both appear on this page alongside your people. Reading what people said in an incident channel is its own ability, **Incident Transcripts**, rather than part of Incidents. Reading an incident and reading every message in its channel are different asks, so granting the first never grants the second. It also needs the workspace to have turned transcript access on, which is a separate admin decision. See [Incident conversations](https://firefight.app/docs/workspace/incident-conversations.md). A few things cannot be granted to anyone: integrations, API keys, this Permissions page, and workspace settings. Those are the controls that decide access itself, so they stay with admins and owners. The number next to each person is how many grants they hold, not how much they can do. Admins show as **admin** because their access does not come from grants. ## When something is not allowed Firefight checks permissions the same way wherever the request comes from, so a person is told the same thing in Slack as an agent is told over the API. In Slack, an action nobody is allowed to take comes back as a message only you can see, naming what was refused and telling you to ask an admin. Nothing happens to the incident. On the dashboard you are sent back to the page you were on with the same message, and a settings screen you cannot change never shows the controls in the first place. Over the API and MCP the caller gets a permission error naming the access it lacks. Some actions are allowed but need a second person to say yes first. Those do not fail, they wait. Which ones, who decides and where they are asked is set in the **Approval rules** block on this page, described in [Approvals](https://firefight.app/docs/gateway/approvals.md). ## Granting Pick who you are granting to, then choose **Grant**. You are granting one of two things. **A single ability** is one capability, such as reading databases in PlanetScale. Search for it by name. Abilities are grouped by the connection that provides them. **A permission set** is a named bundle of abilities granted in one step, covered below. Either way, you then choose which environments the grant covers. Leave every environment unticked and the grant covers all of them. ## Environments If a connection holds separate credentials per environment, a grant scoped to an environment only works there. This is how you let different people and agents reach different environments through the same integration. Grant an AI agent's service key the database abilities scoped to Development, and grant your team the same abilities with no environment restriction. Both use the same capability, and each reaches only what you allowed. For this to work the connection needs credentials wired for that environment. A grant scoped to Production does nothing if the connection has no Production credentials, so set both up together. See [Environments](https://firefight.app/docs/api/integrations#environments). ## Permission sets Granting fifteen abilities one at a time does not scale past a couple of people. A permission set bundles abilities under a name and grants them together. Create one with the **+** next to **Permission sets**, name it for what it lets someone do, such as Database read-only. Then tick the abilities it covers. **Select all** on a group takes every ability from one connection at once. Grant the set the same way you grant a single ability, choosing environments as usual. The set says what someone can do, and the grant says where. :::caution Changing a set changes what everyone holding it can do, straight away. Deleting a set removes it from everyone it was granted to. ::: This is why one set granted twice beats two nearly identical sets. Grant Database read-only to an agent scoped to Development and to your team unrestricted, and when the tool adds a capability you only update the set once. :::note Permission sets are separate from the incident roles under **Configure → Incident Roles**. Incident roles say who is accountable for what during an incident, such as Incident Lead. Being given one grants no access. See [Incident roles](https://firefight.app/docs/incidents/incident-roles.md). ::: ## Changing and removing access Open the person or key to see everything they hold. Each grant shows the ability or set, whether it reads or writes, and which environments it covers. Click the environment button to change the scope without regranting. The bin icon revokes the grant. Revoking takes effect immediately. --- Source: https://firefight.app/docs/getting-started/your-first-incident # Your first incident This guide walks through the full life of an incident in Slack. You'll declare it, coordinate the response, and close it out. It assumes Firefight is already installed in your Slack workspace. ## Try it with a test incident Right after installing, the welcome message in `#incidents` and the dashboard both offer **Declare a test incident**. A test incident gets a channel, an announcement and a postmortem like any other, and is not counted in your metrics. Firefight posts the next step in its channel as you go, so you never have to guess what to do. ## Declare the incident From any channel, type ``` /ff new ``` `/ff` is the short alias for `/firefight` and the two are interchangeable. You can also use the **Create an incident** shortcut from Slack's shortcuts menu or the **Declare incident** button on the dashboard. A dialog opens asking for the essentials. - **Name.** A short description of what's happening, like "Checkout returning 500s". - **Severity.** How bad it is. Out of the box you choose from Critical, Major, or Minor. - **Type.** What kind of incident it is, like Production or Security. Your workspace may ask for more fields here, since admins can customize the declare form. When you submit, Firefight creates a dedicated incident channel named after the date and your incident name, something like `#inc-2026-07-20-checkout-returning-500s`. It announces the incident to your team and posts a welcome message in the channel with the incident's details and quick actions. Everything from here on happens in that channel. Anyone who wants to follow along without joining the channel can click **Subscribe** on the announcement. Firefight then sends them every update it posts about the incident as a direct message. See [Follow an incident without joining it](https://firefight.app/docs/incidents/running-from-the-dashboard#follow-an-incident-without-joining-it). :::tip Don't hesitate to declare. An incident that turns out to be nothing can be closed as a false positive in seconds. That's what the Canceled outcome is for. The expensive mistake is the real incident nobody declared. ::: ## Coordinate the response Inside the incident channel, the `/ff` commands manage the incident. These are the ones you'll use most. **Assign a lead.** Every incident needs one person who owns coordination. ``` /ff lead ``` Run `/ff roles` instead to fill every [incident role](https://firefight.app/docs/incidents/incident-roles.md) at once, lead included. **Post updates.** As the situation develops, keep the record current. ``` /ff update ``` This opens a dialog that starts with the update itself, what is happening and what you are doing next, then lets you change the status and severity alongside it, for example moving from Investigating to Identified once you know the cause. You can also say when the next update is due, and Firefight nudges the lead in the channel when that time arrives, so a quiet incident does not drift. Updates are announced and added to the incident timeline, so people joining late can catch up without scrolling. **Bring in help.** Invite teammates to the channel. ``` /ff invite ``` If you need people urgently, `/ff escalate` notifies them and asks for an acknowledgement, so you know whether help is actually on the way. **Track action items.** When someone says "we should restart the worker pool", capture it before it's lost. ``` /ff action ``` Or react to their message with 💥 (`:boom:`) and Firefight turns that message into an action item. Actions have an owner and a status, and `/ff actions` lists where everything stands. For things that should happen after the incident, like "add an alert for this", use `/ff followup` or react with ▶️ (`:arrow_forward:`) instead. **Catch up.** Joining an incident that's been running for an hour? `/ff timeline` shows the key events so far. ## Resolve it When the issue is fixed and stable, click **Resolve** on the pinned message at the top of the incident channel, or type ``` /ff resolve ``` `/ff close` does the same thing. A dialog asks for closing details, then Firefight marks the incident resolved and announces it. The resolution message offers a **Write the postmortem** button. If the problem comes back, `/ff reopen` picks up right where you left off, with the same channel and the same history. ## After the incident Two things are worth doing while it's fresh. - **Draft the postmortem.** Click **Write the postmortem** on the resolution message, or run `/ff postmortem`. Firefight drafts it from the timeline and the channel messages, ready to edit in the web dashboard. Anything it was not told is left for you to fill in, never made up. - **Say thanks.** `/ff shoutout`, or reacting to someone's message with ❤️‍🔥 (`:heart_on_fire:`), recognizes a teammate who came through. The incident, its timeline, actions, follow-ups, and postmortem all live on in the web dashboard, searchable whenever you need them. --- Source: https://firefight.app/docs/incidents/concepts # Statuses, severities & types Every incident in Firefight carries four pieces of classification. Two of them, **status** and **severity**, change over the incident's life. The other two, **type** and the underlying **lifecycle stage**, describe what kind of thing it is and where it stands. Understanding how they fit together makes everything else in Firefight predictable. ## Lifecycle stages Under the hood, every incident is in exactly one of four stages. | Stage | Meaning | |---|---| | **Triage** | Something's been reported, but it's not yet confirmed as a real incident. | | **Active** | A confirmed incident with an ongoing response. | | **Closed** | The incident is over and was real. | | **Canceled** | It turned out not to be an incident. A false positive or duplicate. | You can't add or rename stages, and you'll rarely deal with them directly. They exist to give your statuses meaning. Every status maps to one stage, so no matter how you customize the labels, Firefight always knows which incidents are live, which are closed, and which never turned out to be real. Closed and Canceled both mean the response is over, so the incident channel's **Resolve**, **Escalate** and **Make me Lead** buttons go away once an incident reaches either one. ## Statuses Statuses are the labels your team actually moves incidents through, and each status belongs to one of the stages above. A new workspace starts with these. | Status | Stage | Typically means | |---|---|---| | Triaging | Triage | Confirming whether this is a real incident | | Investigating | Active | Root cause under active investigation | | Identified | Active | Root cause found, fix in progress | | Monitoring | Active | Fix deployed, watching for stability | | Resolved | Closed | Fully resolved | | Canceled | Canceled | False positive, duplicate, or invalid | You change status with `/ff status` or `/ff update` in the incident channel, and each change is announced and recorded on the timeline. Because each status maps to a stage, moving an incident from Monitoring to Resolved is also what closes it. There's no separate bookkeeping step. Cancelling is its own action rather than just another status change. Use `/ff cancel`, or the **Cancel incident** button that sits next to **Accept incident** while an incident is still in triage. A cancelled incident keeps its channel and timeline, but it never counts as resolved, so it stays out of your time-to-resolve figures and is not offered a postmortem. If you cancel something and then find out it was real, reopen it with `/ff reopen` and carry on in the same channel. If you add a second status to the Closed or Canceled stage, responders are asked which one applies when they resolve or cancel. That is the simplest way to record why something was cancelled. Add "Duplicate" and "False positive" alongside "Canceled" and every cancellation says which it was, with no extra field to fill in. While a stage holds a single status there is nothing to choose, so nobody is asked. Statuses are fully customizable in **Settings → Statuses**. Rename them, reorder them, add your own within any stage, and choose which one new incidents start in. ## Severities Severity says how bad the incident is, independent of where it is in its lifecycle. | Severity | Typically means | |---|---| | **Critical** | Service-wide outage or data loss | | **Major** | Significant feature degradation | | **Minor** | Limited impact or a workaround exists | Severity is set when the incident is declared and can be changed at any time with `/ff severity` as you learn more. It's normal for an incident to be declared Minor and upgraded once the blast radius is clear. Severity drives visibility. It appears in announcements, channel headers, and dashboard filters, and alert routing rules can set it automatically for incidents created from alerts. Customize the list and its order in **Settings → Severities**. ## Types Type answers "what kind of incident is this?" and is useful for filtering and for spotting patterns across incidents. The defaults are **Production**, **Security**, **Infrastructure**, **Data**, and **Third Party**. Type is chosen at declare time and is mostly an organizational tool. Reviewing last quarter's incidents by type tells you where your reliability effort should go. Customize in **Settings → Types**. ## The timeline Every incident page in the dashboard has a timeline: the full story of what happened, in order, with who did it and what it was about. Each entry names its subject rather than just the kind of thing that happened. - **Status, severity, lead, and field changes** show the before and after, one row per field. When the summary is rewritten, the card shows the new text. Click **Show previous version** to read what it replaced. - **Runbooks** name the runbook, and the name opens it. A runbook that attached itself also says which rule matched, such as "Severity is one of Critical, Major". - **Action items** show the item's description, its status, and who holds it. Clicking one highlights it in the Actions panel. - **Escalations** name the person, with their avatar, and the reason given. Reminders and acknowledgements sit alongside. - **Pinned messages** quote the message and link to it in Slack. - **Alerts** name the alert that attached or resolved. **Related, duplicate, and merged incidents** name the other incident and link to it. - **Roles** name the role and the person. Entries performed by a person carry their avatar. Entries Firefight performed on its own, such as attaching a runbook by rule or sending an escalation reminder, are attributed to **Firefight**. On the incidents list, click anywhere on a row to open the incident. The ID and name are links as well, so an incident declared without a name is still one click away. The incident page is also where you change any of this. See [Running an incident from the dashboard](https://firefight.app/docs/incidents/running-from-the-dashboard.md). ## Timeline notes The timeline above records what Firefight did. It does not record what the team worked out. The theory someone floated at 14:20 and the rollback that finally worked live in the channel, and they go with it. When an incident is resolved or cancelled, Firefight reads the channel once and adds the milestones of the investigation to the timeline as notes. A note is one line in the past tense, attributed to the person who said it and placed at the time they said it, so it sits inside the conversation and reads alongside the rest of the timeline. Each one is labelled **AI-noted**, quotes the message it came from, and links straight to that message in Slack. Firefight notes eight kinds of moment. | Kind | Example | |---|---| | Hypothesis | Diego suspected the 14:02 deploy | | Finding | Uros confirmed the connection pool was exhausted on replica 2 | | Root cause | Diego identified the migration lock as the root cause | | Mitigation | Uros rolled back the 14:02 deploy | | Decision | The team decided to fail over rather than wait for the vendor | | Blocker | The team was blocked on Datadog support with no ETA | | Impact | Checkout was down for EU customers only | | Recovery | Error rate returned to baseline | Greetings, acknowledgements, unanswered questions, and anything already on the timeline are left out. Status changes and escalations are recorded as events in their own right, so a note never repeats one. Nothing is posted to the channel. Notes appear on the incident's timeline in the dashboard, and they flow into the postmortem draft, `/ff catchup` on a closed incident, the API, and anything reading your timeline over [MCP](https://firefight.app/docs/api/mcp-server.md), so an agent can ask what the root cause of a past incident was without reading a single message. ### Dismissing a note A note is a reading of a conversation, and a reading can be wrong. If one credits the wrong person or takes a joke for a decision, open the note's actions menu on the timeline and choose **Dismiss note**. Dismissed notes are not deleted. They collect at the end of their day on the timeline under "N dismissed notes", which you can expand to see what was dismissed and by whom, so the correction stays visible. Everywhere else a dismissed note is gone. It leaves the API, MCP, `/ff catchup`, and any postmortem drafted afterwards. An incident is read once, however long it ran and however many messages it holds. See [AI in Firefight](https://firefight.app/docs/ai/overview.md) for what the AI sees and [AI data handling](https://firefight.app/docs/ai/data-handling.md) for what is stored and redacted. ## How they work together A typical incident reads like this. Someone declares a **Production** incident at **Major** severity, and it starts in the **Investigating** status. When the scope becomes clear, severity is raised to **Critical**. The response moves from **Investigating** to **Identified**, then to **Monitoring**, and ends at **Resolved**, which also closes the incident. A practical rule of thumb is that **status describes the response and severity describes the impact**. If you need a label for waiting on a vendor, add a status. If you need a label for a full outage, add a severity. --- Source: https://firefight.app/docs/incidents/incident-roles # Incident roles An incident role says who is handling a particular part of the response, like the Incident Lead who coordinates it. You assign roles inside the incident channel and change them as the incident develops. Firefight comes with one role. You define the rest. | Role | Description | | --- | --- | | Incident Lead | Coordinates incident response and makes decisions | Incident Lead is built into Firefight. It is there from your first incident and you cannot delete or disable it, because incident lists and the incident page always show who is leading. ## What a role means A role names one accountable person, not everyone doing the work. Making someone the Incident Lead does not mean they fix the incident alone, it means they own coordinating it and saying where things stand. The same goes for any role you define. Work that several people share belongs in action items and follow-ups instead, which any number of people can pick up. Assigning a role does not change what someone can do in Firefight. That is set by their [access level](https://firefight.app/docs/workspace/members-and-access.md). ## Defining your own roles **Configure → Incident Roles** lists every incident role. Click **Add Role** to define your own, for example Operations Lead or Communications Lead, with a name and a description of what the role is responsible for. Drag the rows to set the order they appear in when assigning. To stop offering a role, disable it. Past incidents keep their assignments, and you can re-enable it whenever you want it back. Delete is only for a role no incident has ever used. ## Assigning roles during an incident In the incident channel, run: ``` /ff roles ``` An **Incident Roles** dialog opens with one picker per role, each already showing whoever holds it. Change any of them and save. Firefight posts a single message in the channel summarising what changed. To assign only the lead, `/ff lead` still opens a dialog for that one role. Each role holds one person at a time, so choosing someone new hands the role over rather than adding a second holder. Clearing a picker and saving leaves the role unassigned. The Incident Lead is the exception. Once an incident has a lead it keeps one, so you can hand the lead to someone else but not leave it empty. Roles stop changing once an incident is resolved or canceled. Every change announces itself in the incident channel, and that channel is on its way to being archived, so Firefight says which role can no longer be changed rather than posting into a room nobody is reading. Reopen the incident if the response is genuinely starting again. ## Where roles show up The lead appears on the incident page and in every incident list across the web dashboard. Anyone catching up can see who is coordinating without opening the incident. Every other role you have assigned appears in a Roles panel on the incident page, each one next to the person holding it. Roles nobody holds are left out, so the panel shows the roster as it actually stands. Every change is recorded on the incident timeline, alongside status changes and updates. When you look back at the incident, you can see who took a role and when they took it. If you forward incidents to your own systems, role changes arrive as [webhook](https://firefight.app/docs/api/webhooks.md) events. The lead sends `lead.assigned`. Every other role sends `role.assigned` when someone takes it and `role.unassigned` when it is cleared. ## Working with roles from an AI agent An AI agent connected to Firefight can handle the roster as well as a person can, which is useful when the agent is already triaging the incident. `get_incident` returns every role your workspace has defined together with whoever currently holds it, so the agent knows the roster before it changes anything. `assign_incident_role` then puts one person in one role, or clears it. The same rules apply as in Slack. One person per role, and the lead can be handed over but not left empty. See [Connect AI agents](https://firefight.app/docs/api/mcp-server.md). :::note Assign a lead early, even for small incidents. A named coordinator is the single highest-leverage habit in incident response, and it also gives AI-generated [postmortems](https://firefight.app/docs/postmortems/overview.md) a clear account of who was driving. ::: --- Source: https://firefight.app/docs/incidents/runbooks # Runbooks A runbook is your team's answer to "we've seen this before, here's what to do." In Firefight, runbooks are structured objects rather than wiki pages: a short summary, a rich procedure written in Markdown, and an ordered list of steps. When an incident matches a runbook's conditions, the runbook attaches itself and appears in the incident channel, so the procedure finds the responder instead of the responder hunting for the procedure. Runbooks live at **Settings → Runbooks** and require workspace admin access to manage. ## Anatomy of a runbook | Part | What it is for | |---|---| | Name | What the procedure is called, like "Database failover" | | Summary | A short description of when to use it. Shown in listings, in Slack, and to AI agents deciding whether the runbook is relevant | | Content | The full procedure, edited with formatting (headings, lists, code blocks) and stored as Markdown | | External URL | An optional link to a page elsewhere, like a wiki article or a dashboard, if part of the procedure lives outside Firefight | | Steps | An ordered list of discrete actions responders should take, each with a title and instructions | Content and steps play different roles. Content is the narrative: context, decision points, queries to run, things to check. Steps are the checklist: concrete actions that responders can pick up, complete, and track one by one. ## Conditions and automatic attachment Each runbook carries conditions that decide which incidents it applies to. You can match on: - **Incident type**, like Production or Security - **Severity**, like Critical or Major - **Custom fields**, including catalog references, so a runbook can target incidents affecting a specific service, environment, or any other field you have defined Every condition supports "is one of" and "is not one of", and an incident must match all conditions for the runbook to attach. A runbook with no conditions attaches to every incident. Conditions are evaluated when an incident is declared, and again whenever the incident is updated. That second pass matters: severity and type are usually set at declaration, but the affected service often gets filled in later. A runbook scoped to a service attaches the moment that field is set. Attachment only ever adds. Updating an incident never detaches a runbook that no longer matches, so a procedure that was relevant stays visible. Every automatic attachment is recorded on the incident timeline with the rule that matched, for example "Matched Severity is one of Critical, Major", so a responder can see why a procedure appeared. A runbook attached by hand records who attached it instead. A runbook with no conditions attaches to nothing on its own. If you want one on every incident, turn on **Attach to every incident** and the Conditions column says so. Leaving both off is a valid choice: the runbook waits until someone attaches it by hand. ## Attaching one by hand Not every procedure is worth a condition. Three ways to pull one in when you need it: - `/ff runbook` in the incident channel, or **Attach runbook** in `/ff home`, which opens a picker of everything not already attached - The **Runbooks** panel on the incident page in the dashboard - The `attach_runbook` [MCP tool](https://firefight.app/docs/api/mcp-server.md), so an agent can bring in the procedure it thinks applies All three post the runbook in the incident channel exactly as an automatic attachment does. Attaching one twice does nothing the second time. ## In the incident channel When a runbook attaches, Firefight posts a message in the incident channel listing every step, and that message is the checklist you work from. Each step carries an **I can take this** button. Press it and the step becomes an [action item](https://firefight.app/docs/incidents/slack-commands.md) assigned to you, and the row turns into **Mark as done**. Press that and the row is struck through with your name on it. The message updates in place, so a runbook never fills the channel with posts no matter how many steps it has. Nothing is created until someone takes a step, which means the incident's action list reflects the work people actually did rather than every step in the procedure. **View runbook** opens the full procedure: every step with its complete instructions, plus a person picker on each one. That is where an incident lead hands work out. Assigning a step to someone creates or moves its action and posts it in the channel naming them, with the controls to work it, so they never have to come back to the runbook message. Each completed step posts a short line naming the runbook it came from, so progress through a procedure is visible without scrolling back. `runbook.attached` is recorded on the incident timeline when a runbook attaches, and each step you take records the usual `action.created`, `action.reassigned`, and `action.completed` events. All are available as [outbound webhook events](https://firefight.app/docs/api/webhooks.md). ## For AI agents and the API Runbooks are first-class context for AI. Attached runbooks are included in the incident context that Firefight's own [incident responder](https://firefight.app/docs/ai/incident-responder.md) sees, and any external agent connected over [MCP](https://firefight.app/docs/api/mcp-server.md) can browse them with two read-only tools: `search_runbooks` to scan names and summaries, and `get_runbook` to pull the full procedure and steps. The [REST API](https://firefight.app/docs/api/overview.md) covers the same ground at `/api/v1/runbooks`, addressable by ID or slug, under the `runbooks` permission resource. It reads with `GET` and writes with `POST`, `PATCH` and `DELETE`, and the `upsert_runbook` MCP tool does the same over MCP. Steps and conditions are only touched when you send them, so changing a summary never silently clears the procedure. Deleting is refused while a runbook is attached to an incident, and the refusal says how many. ## Writing runbooks that work A few habits make runbooks dramatically more useful, for humans and agents alike: - Lead with symptoms. The first thing a responder needs to confirm is "is this what is actually happening?" - Make each step a single imperative action, not a paragraph. "Check replication lag on the standby" beats a wall of prose - Include the exact queries, dashboard links, and commands. Every lookup a responder has to improvise costs minutes - Keep the summary honest and specific. It is how the right runbook gets found, by people scanning a list and by agents choosing what to read - Review after every incident that used the runbook. Stale procedures are worse than none --- Source: https://firefight.app/docs/incidents/running-from-the-dashboard # Running an incident from the dashboard The dashboard is not a read-only view of what happened in Slack. You can declare an incident from the incidents list, and run it from its own page, using the same dialogs Firefight asks for in Slack. Which surface you use is a matter of where you already are. On an incident's page the controls sit in two places. Things that are a property of the incident are edited where they are shown, so you click the thing you want to change. Things that move the incident from one state to another live in the **⋮** menu at the top right, next to the channel link. ## Declaring an incident **Declare incident** at the top of the incidents list opens your workspace's Declare form, the same one `/ff new` opens in Slack, so it asks for the same fields including any custom ones. Firefight creates the incident channel straight after, which takes a moment. Until it exists, the channel button on the incident's page is greyed out and shows the name the channel is about to get. If it stays greyed out, creating the channel failed. ## Change severity or status Click the severity or the status badge at the top of the page. Both open the **Post an update** dialog with the incident's current values filled in. They open the whole dialog rather than a short list because that is what your workspace configured. If the Update form asks for a written message, or for a custom field, it asks for it here too. That is the same dialog `/ff update`, `/ff status` and `/ff severity` open in Slack, so an update posted from the dashboard reads the same in the channel and records the same timeline entry. ## Assign the lead and the other roles Click the person under **Lead**, or the person next to any role in the Roles panel. Both are dropdowns of everyone in the workspace. Every configured role is listed whether or not anyone holds it, so a role nobody has taken yet can be filled here. Picking **Unassigned** clears a role. The incident lead is the one exception and cannot be cleared, only handed to someone else, so the incident always has someone accountable. Each change announces itself in the incident channel exactly as `/ff lead` and `/ff roles` do. ## Resolve, cancel and reopen Open the **⋮** menu. | Action | What happens | |---|---| | **Post an update** | The Update dialog, the same as clicking a badge. | | **Resolve incident** | The Resolve dialog. Closes the incident, which is what makes a postmortem available. | | **Cancel incident** | The Cancel dialog. The incident stays with its channel and timeline, but never counts as resolved. | | **Reopen incident** | Only on an incident that is already over. Puts it back on your default live status. | Resolve and Cancel ask what your workspace's Resolve and Cancel forms ask, which may include custom fields. If you have more than one status in the Closed or Canceled stage, they ask which one applies. If you have only one, they do not, because there is nothing to choose. See [Statuses, severities & types](https://firefight.app/docs/incidents/concepts.md). An incident that is over stops offering the badges and the role pickers, because every one of those changes announces itself in a channel that may already be archived. Reopen it first. It also stops taking new actions, and its runbook steps can no longer be claimed, because both are work during the incident. Follow-ups can still be added, since they are the work that comes after. ## Work the action items Every action and follow-up in the sidebar has a menu on hover. - **Pick up** takes an unclaimed item and marks it in progress. It appears only while nobody holds it. - **Mark done** completes it. - **Assign to** hands it to anyone in the workspace. Taking an item yourself and handing it to someone else are different things. Taking your own work is silent, because you already know. A handover posts in the incident channel, so the person finds out. ## Claim a runbook step Expand a runbook in the Runbooks panel to see its steps, who is on each one, and which are done. **Claim** on a step takes it, which creates the action item behind that step and puts it in the sidebar alongside everything else. Claiming a step somebody else already holds hands it over rather than creating a second copy. ## Link incidents together The **⋮** menu also carries the two relationships. **Link to an incident** records that two incidents are related. Both name the other on their timelines and neither changes status. **Mark as duplicate** is for when this incident turns out to be the same event as another. This one is cancelled and points at the one you pick, which stays open as the real incident. Both match `/ff link` and `/ff duplicate` in Slack. ## Bring people in The **⋮** menu carries three ways to involve other people, all of which post in the incident channel. **Ask someone to pick this up** names one person and says why you need them. They get a message of their own with an acknowledge button, the ask goes in the incident channel, and Firefight reminds them if they do not answer. Reach for this when you need a particular person to respond, not just to watch. **Bring people in** adds responders to the incident channel so they can follow along and join in. It asks nobody for anything, which is the difference from the one above. **Give a shoutout** thanks someone for their work, posted in the incident channel where everyone working the incident sees it. All three match `/ff escalate`, `/ff invite` and `/ff shoutout` in Slack. Each offers the people already in your workspace. To pull in somebody Firefight has not seen yet, invite them from Slack. An incident that has been resolved or cancelled shows these greyed out with the reason, as does one whose channel Firefight is still creating, because all three need a channel to post in. ## Follow an incident without joining it Not everyone who cares about an incident belongs in its channel. A manager, someone in support, the on-call for a neighbouring service. The **Subscribers** card on the incident page lists who is following it, and **Subscribe** adds you. From then on, every update Firefight posts about the incident reaches you as a direct message from Firefight, the same message it posts in the announcement thread in your incidents channel. That covers status, severity and lead changes, written updates, escalations, and the incident being resolved or reopened. Actions, follow-ups and channel conversation are not included, since those belong to the people working the incident. Each message names the incident at the top and ends with **Open channel**, **Incident homepage** and **Unsubscribe**, so a message about one of several incidents you follow is never in doubt. **Unsubscribe** on the card stops the messages, and so does the button at the foot of any message. The **Subscribe** button on the announcement in your incidents channel subscribes you from Slack. Firefight confirms with a note only you can see, which carries **Unsubscribe** in case you clicked by mistake. ## Open the channel The channel name at the top right opens the incident's channel in your Slack app. ## What still lives in Slack `/ff catchup` posts a written catch-up into the channel, and has no dashboard equivalent. See the [Slack command reference](https://firefight.app/docs/incidents/slack-commands.md) for everything available there. --- Source: https://firefight.app/docs/incidents/slack-commands # Slack command reference Firefight registers two identical slash commands, `/firefight` and its short alias `/ff`. These docs use `/ff`. Most commands act on the incident whose channel you run them in, so unless noted otherwise, run them **inside an incident channel**. Most of what these commands do can also be done from the incident's page in the dashboard, using the same dialogs. See [Running an incident from the dashboard](https://firefight.app/docs/incidents/running-from-the-dashboard.md). ## Work from anywhere | Command | What it does | |---|---| | `/ff new` | Declare a new incident. Opens the declare dialog, then creates the incident channel. | | `/ff list` | List the incidents that are currently live. | | `/ff home` | Open the Firefight home dialog, an overview with quick actions for common tasks. | ## Run the response | Command | What it does | |---|---| | `/ff lead` | Assign the incident lead. | | `/ff roles` | Assign every incident role in one dialog, one person each. | | `/ff status` | Change the incident's status. | | `/ff severity` | Change the incident's severity. | | `/ff update` | Post an incident update. One dialog covers status, severity, a written summary of where things stand, and when the next update is due. | | `/ff summary` | Edit the incident's summary. | | `/ff invite` | Invite teammates into the incident channel. | | `/ff escalate` | Urgently notify people and request an acknowledgement, so you know help is coming. | ## Track the work | Command | What it does | |---|---| | `/ff action` | Create an action item, something to do now as part of the response. | | `/ff actions` | List the incident's action items and their status. | | `/ff followup` | Create a follow-up, something to do after the incident is over. | | `/ff followups` | List the incident's follow-ups. | | `/ff runbook` | Attach a runbook to this incident, picking from the ones not already on it. | | `/ff timeline` | Show the incident's timeline of key events. | | `/ff catchup` | Get an AI-written summary of the incident so far, useful when joining late. | Every action item and follow-up is posted in the channel with its own controls. **I can take this** assigns it to you, **Mark as done** closes it out, and the person picker beside them hands it to someone else at any point before it is done. Handing one over posts it again naming the new holder, so they can work it where they were told about it. The same controls sit on every row in `/ff actions` and `/ff followups`, so you can work the whole list from one screen. Completing anything posts a short line saying what was finished and who finished it, linking back to where it came from. Work spread across an incident is easy to miss otherwise, because a finished item only strikes itself through in place and Slack does not announce an edit. ## Connect incidents | Command | What it does | |---|---| | `/ff link` | Link this incident to another one. | | `/ff relate` | Mark another incident as related to this one. | | `/ff duplicate` | Mark this incident as a duplicate of another. | ## Wrap up | Command | What it does | |---|---| | `/ff close` | Close the incident. `/ff resolve` does the same thing. | | `/ff cancel` | Cancel an incident that turned out not to be one. | | `/ff reopen` | Reopen a resolved or canceled incident with the same channel and history. `/ff open` does the same thing. | | `/ff postmortem` | Generate an AI-drafted postmortem from the incident's timeline and discussion, ready to edit in the dashboard. | | `/ff shoutout` | Recognize a teammate who came through during the incident. | ## Reaction shortcuts Inside an incident channel, reacting to any message turns it into a record. No command needed. | Reaction | What it does | |---|---| | 💥 `:boom:` | Turn the message into an **action item**. | | ▶️ `:arrow_forward:` | Turn the message into a **follow-up**. | | ❤️‍🔥 `:heart_on_fire:` | Give the message's author a **shoutout**. | ## Mention the bot In an incident channel, mention **@Firefight** in a message to ask questions about the incident. It answers from the incident's own timeline and discussion. --- Source: https://firefight.app/docs/postmortems/editing-and-revisions # Editing & revisions Every postmortem opens in a full-page rich text editor in the web dashboard. Anyone in your workspace can edit it until it is marked completed. If you have not created one yet, see [Postmortems](https://firefight.app/docs/postmortems/overview.md) for where they live and [Generating with AI](https://firefight.app/docs/postmortems/generating-with-ai.md) for drafting. ## The editor The editor behaves like a modern document editor. Select any text and a floating toolbar appears with bold, italic, underline, strikethrough, and inline code. The document supports headings, lists, task lists with checkboxes, blockquotes, and code blocks. Changes save automatically. Shortly after you stop typing, the header shows a brief **Saved** confirmation, so there is no save button to remember. Saving also happens when you navigate away, so a quick edit on the way out is not lost. ## If somebody else changes it while you are writing A postmortem can also be written by an agent or a script over the [API](https://firefight.app/docs/api/using-the-api#postmortems). If one of them replaces the document while you have it open, your next save is refused rather than overwriting what they wrote, and a message appears above the editor saying so. Saving stops at that point, so nothing you type afterwards is written. Copy anything you want to keep, then choose **Reload** to see their version and paste it back in. Their write is on the [revision history](#revision-history) too, so nothing is lost either way. Selecting text also exposes the AI rewrite button. See [Rewrite a section with AI](https://firefight.app/docs/postmortems/generating-with-ai.md) for how that works. ## Changing status The status badge in the header shows where the document stands. Use the menu in the top-right corner of the postmortem page: - **Mark as Completed** locks in the postmortem as finished. - **Reopen as Draft** appears on a completed postmortem and puts it back into editing. The same menu offers **Export as PDF** and **Export as Markdown** if you need to share the document outside Firefight. ## Revision history Every meaningful change to the postmortem is captured as a revision. That includes the original AI generation, manual edits, and AI rewrites. Each revision records who made the change, what kind of change it was, when it happened, and a snapshot of the document at that point. Click **Revisions** in the header to open the history panel. Revisions are listed newest first with the editor's name and the change type: | Change type | What it means | | --- | --- | | Generated by AI | The initial AI-written draft | | Edited | A manual edit in the editor | | AI rewritten | A passage was rewritten through the AI rewrite dialog | Select a revision to preview it. The preview highlights what changed between that revision and the next one, with additions and deletions marked inline. If an earlier version was better, click **Restore this version** and it becomes the current document. Restoring is itself recorded as an edit, so nothing in the history is ever lost. :::tip Restore is safe to experiment with. Because the version you replace is captured as a new revision, you can always get back to it from the same panel. ::: --- Source: https://firefight.app/docs/postmortems/generating-with-ai # Generating with AI Firefight can write the first draft of a postmortem for you. The AI reads the incident record and produces a structured document in the [nine standard sections](https://firefight.app/docs/postmortems/overview.md), written in a blameless tone that focuses on systems and processes rather than individuals. ## Generate from Slack When an incident is resolved, the resolution message in its channel offers a **Write the postmortem** button. You can also run `/ff postmortem` in the channel. Either works only once the incident is closed and while it has no postmortem yet, and takes under a minute. ## Generate from the dashboard On the incident page in the web dashboard, the postmortem card appears once the incident is closed. Click **Generate draft**. You land on the postmortem page, which shows a placeholder while the draft is written and switches to the editor as soon as it is ready. ## What the draft is built from The AI works only from what Firefight recorded about the incident. That includes: | Input | Examples | | --- | --- | | Incident details | Name, severity, status, who declared it, who led it, key timestamps, duration | | Custom fields | Any values your team filled in on the incident forms | | Timeline | The recorded events from declaration through close | | Channel discussion | A narrative summary of the conversation in the incident channel | | Actions and follow-ups | What was tried during response and what remains open | | Shoutouts | Recognition given during the incident | This is why disciplined use of `/ff update`, actions, and the incident channel pays off. The richer the record, the better the draft. ## Sections the record cannot fill The AI writes only what the timeline and the channel messages support. It never guesses a cause, an impact or a fix. Every heading still appears. The Timeline is always filled in, and under any heading it had nothing for you find a short note asking you to add what you know. Post in the incident channel while it runs and the draft has more to work from. ## If generation fails If the draft cannot be written, the postmortem page says so and offers **Try again** and **Start blank**. The person who asked for the draft is also told in the incident channel. ## Generate over the API or MCP `POST /api/v1/incidents/:id/postmortem` with `generate: true` starts the same draft, and `start_postmortem` does it over MCP. Drafting takes a moment, so the response comes back straight away with a `generation_state`. Read the postmortem back until that state clears, then edit it like any other. Leave `generate` out and you get the blank one described below. ## Start blank instead If you prefer to write from scratch, click **Start blank** on the incident page card. You get an empty document in Draft status with the incident's identifier and name as the title. Starting blank does not use AI. ## Rewrite a section with AI Inside the editor you can point AI at any passage. Select some text and click the sparkles button in the floating toolbar. A **Rewrite with AI** dialog opens where you describe what you want, for example "make this more concise" or "expand on the root cause". The AI has access to the incident context, including the timeline, summary, and actions, so it can add real detail rather than padding. The rewritten text replaces your selection, and you can undo it like any other edit. ## The AI drafts, humans own accuracy Treat the generated document as a strong first draft, not a finished artifact. The AI only knows what made it into the incident record. It cannot know about the hallway conversation, the dashboard someone eyeballed, or the context behind a decision. Before you mark a postmortem completed, have someone who was on the incident verify the facts, sharpen the contributing factors, and make sure every action item has an owner. :::tip Review the Summary and Key contributing factors sections most carefully. They are the parts future readers rely on, and the parts where missing context does the most damage. ::: --- Source: https://firefight.app/docs/postmortems/overview # Postmortems A postmortem is the written record of an incident after it is over. It captures what happened, why it happened, what the impact was, and what your team changes so it does not happen again. In Firefight every closed incident can have exactly one postmortem, and it lives alongside the incident's timeline, actions, and follow-ups. You can let AI write the first draft from the incident record, or start from a blank page and write it yourself. Either way the document stays editable by your team until you mark it completed. See [Generating with AI](https://firefight.app/docs/postmortems/generating-with-ai.md) for the drafting flow and [Editing & revisions](https://firefight.app/docs/postmortems/editing-and-revisions.md) for the editor. ## The nine sections A generated postmortem is structured into nine sections, in this order: | Section | What it covers | | --- | --- | | Summary | The problem, impact, causes, and steps to resolve at a glance | | Introduction | Context for readers who were not in the room | | Timeline | Every recorded event, from declaration to resolution, in order | | Deeper dive | A detailed narrative of how the incident unfolded | | Impact | Who and what was affected, and for how long | | Resolution | How the incident was brought to an end | | Key contributing factors | The conditions that allowed the incident to happen | | What went well | Response practices worth keeping | | Action items | Concrete follow-ups that prevent a repeat | Every heading is always present on a generated document. The Timeline is never empty, and any section the AI had nothing for carries a short note asking you to add what you know. See [Generating with AI](https://firefight.app/docs/postmortems/generating-with-ai#sections-the-record-cannot-fill). The sections are a starting structure, not a cage. The document is free-form rich text, so you can expand, trim, or reorganize sections as your team sees fit. ## Statuses A postmortem moves through four statuses. The current status is always visible as a badge on the postmortem page and on the incident page card. | Status | Meaning | | --- | --- | | Draft | The document is being written and edited | | In progress | Someone is actively working on it, set over the API or MCP | | In review | The document is being reviewed by the team | | Completed | The postmortem is finished | A freshly generated or blank postmortem starts in Draft. While AI generation runs, the page shows a placeholder until the draft arrives, and the status stays Draft throughout. When the write-up is done, use the menu in the top-right of the postmortem page to mark it as Completed. If something needs another pass, the same menu reopens it as a Draft. ## Over the API and MCP Everything above can be done without opening the dashboard. `GET`, `POST` and `PATCH` on `/api/v1/incidents/:id/postmortem` read the write-up, open it, and change its body or status, and MCP has `get_postmortem`, `start_postmortem`, `update_postmortem` and `set_postmortem_status` for the same four things. Sending a new body replaces the whole document rather than appending to it, so read the current one first if you mean to add a section. Reading gives you a version, and sending a body means sending that version back, so a write built on a document somebody has since changed is refused rather than replacing their work. Every version is kept either way. See [Using the API](https://firefight.app/docs/api/using-the-api#postmortems) and the [MCP server](https://firefight.app/docs/api/mcp-server#writing-the-postmortem). ## Where postmortems live Each incident page in the web dashboard has a postmortem card. While the incident is still open, the card reminds you that a postmortem becomes available once the incident is resolved. After the incident is closed, the card offers **Generate draft** and **Start blank**. Once a postmortem exists, the card shows its status and links to the full-page editor. In Slack, the resolution message offers **Write the postmortem** as well. From the postmortem page you can also export the document as PDF or as Markdown through the menu in the top-right corner. :::tip Write the postmortem while memories are fresh. Generating a draft right after closing the incident gives reviewers something concrete to react to, which is much faster than starting from nothing a week later. ::: --- Source: https://firefight.app/docs/workspace/incident-channels # Incident channels Every incident gets its own Slack channel, and once the incident is over that channel is a room nobody is reading. Firefight archives it after a delay you choose under **Settings → Workspace**. The messages stay in Slack and the incident keeps its timeline. The setting is admin-only, like the rest of the Workspace page. ## Choosing the delay **Archive the channel** offers a fixed set of delays, counted from the moment the incident ends: | Choice | What happens | |---|---| | Immediately | Firefight archives the channel as soon as the incident is resolved or cancelled. | | 15 minutes, 1 hour, 6 hours, 24 hours | The channel stays open for a short wrap-up, then archives on its own. 1 hour is the default. | | 3 days, 7 days, 30 days | The channel stays open long enough to write the postmortem in it, then archives. | | Never | Firefight leaves every incident channel open. Archive them yourself in Slack when you want to. | Pick one and press **Save changes**. Choosing **Never** keeps the delay you had, so switching archiving back on later returns to it rather than to the default. ## When it applies The delay starts when an incident is resolved or cancelled, whichever ends it, and when an incident is marked as a duplicate of another, since the duplicate's channel is over too. A change to the setting applies to incidents that end after you save. An incident that was already resolved keeps the delay it was given when it ended. Reopening an incident unarchives its channel, and Firefight skips the archive for an incident reopened before its delay runs out, so a team that has gone back to work in the channel keeps it. --- Source: https://firefight.app/docs/workspace/incident-conversations # Incident conversations Firefight stores what people say in an incident channel so it can summarise the incident, note what was worked out onto the timeline, and draft a postmortem from it. Two settings under **Settings → Workspace** decide who else can read those messages and how long they are kept. Both settings are admin-only. They sit on the workspace rather than on a permission, because turning access on is a decision about your data rather than a decision about one person. ## Letting agents read the conversation **Let AI agents read incident conversations** is off when you start. While it is off, nothing outside Firefight's own AI features can read the messages, and a request for them is refused whoever makes it, admins included. Turning it on does not by itself let anyone read anything. A caller also needs the **Incident Transcripts** ability, which is granted like any other under [Permissions](https://firefight.app/docs/gateway/permissions.md). Both have to be true, so an agent you granted every incident ability still reads nothing until you turn this on, and turning it on gives nothing to an agent you never granted the ability. :::caution The transcript is the rawest data in your workspace. Secrets matching known token formats are redacted before anything is stored, and each message says whether it was, but names, customer details and links are not. See [AI data handling](https://firefight.app/docs/ai/data-handling.md) for exactly what is scanned. ::: ## Reading it There is no transcript viewer on the dashboard. The conversation is already in Slack, where you can scroll it, so this is for agents and scripts that need the reasoning behind a timeline rather than the timeline itself. Over MCP, `get_incident_transcript` takes the incident and returns its messages oldest last, each with who said it, when, the text, and whether anything in it was redacted. Over REST it is `GET /api/v1/incidents/:id/transcript`, which returns the same thing. Both return 100 messages by default and 500 at most. A long conversation comes back from the end, since the end is usually the part worth reading, and each response carries a `more_before` cursor. Pass it back as `before` to walk further into the past. When there is nothing older left, `more_before` comes back empty. ```bash curl "https://app.firefight.app/api/v1/incidents/3f8c1a90-2b7e-4d15-9c04-6ea5b3d71f28/transcript?limit=50" \ -H "Authorization: Bearer ff_4kWm2xPqR8vNcT6yBhJd3fLzGaU9sEnQoXwA" ``` With the setting off, both surfaces refuse and say so, rather than returning an empty conversation that would read as a quiet incident. ## How long conversations are kept **Keep conversations for** takes a number of days counted from when the incident ends, and defaults to 30. Leave it empty to keep conversations for good. The clock starts when the incident is resolved or canceled rather than when it was declared, so a long incident keeps its whole conversation while it runs and for the full window afterwards. Postmortems are usually written the morning after, and generating one reads the conversation, so a window of a few days is enough for that. Clearing the messages does not clear what the team worked out. That survives on the incident's [timeline](https://firefight.app/docs/incidents/concepts#timeline-notes), where each note carries the quote and the person behind it, and in the postmortem where one was written. What you lose is the conversation around those moments. Changing the number applies from the next day's purge onward. Shortening the window is not reversible, since the messages it covers are deleted rather than hidden. --- Source: https://firefight.app/docs/workspace/members-and-access # Members & access Everyone who uses Firefight is a member of your workspace. Each member has an access level that decides whether they can change how the workspace is set up, and it stays the same whatever they are doing on any given incident. ## How people become members There is no invite flow to manage. Anyone in your Slack workspace who uses Firefight becomes a member automatically. Running a `/ff` command, being picked as incident lead, or signing in to the web dashboard with Slack is enough. Firefight creates the membership on the spot using the person's Slack profile. The first person to connect the workspace becomes its owner. Everyone after that joins as a regular member. ## The members list **Settings → Members** shows everyone in your workspace. For each person you see their name, email, and avatar from Slack, their access level, and when they joined. | Access | Meaning | | --- | --- | | Owner | The person who first connected the workspace | | Admin | Can change workspace settings as well as respond to incidents | | Member | Everyone else who uses Firefight | Admins and owners can edit the settings screens. Members respond to incidents with the full set of `/ff` commands, and read the dashboard without being able to reconfigure the workspace. A settings screen a member cannot change shows the same list without the add, edit and reorder controls. An admin can open a settings screen to a specific member without making them an admin, by granting it under **Gateway → Permissions**. Grant someone runbooks, and the **Runbooks** screen gives them the same controls an admin sees, everywhere else stays read-only. Integrations, API keys, Permissions itself, and workspace settings cannot be granted, and stay with admins and owners. [Permissions](https://firefight.app/docs/gateway/permissions.md) and [Activity](https://firefight.app/docs/gateway/activity.md) under **Gateway** are visible to admins and owners only. [Approvals](https://firefight.app/docs/gateway/approvals.md) is visible to everyone, the sidebar shows how many requests are waiting, and approving or denying a request needs the access the request asks for. Activity is the workspace's audit log: every settings change anyone makes on the dashboard is listed there, alongside what agents and API keys did. You never need to raise someone's access to give them a job during an incident. Making a member the Incident Lead leaves their access exactly as it was. See [Incident roles](https://firefight.app/docs/incidents/incident-roles.md). ## Granting more than an access level allows To grant someone reach beyond their access level, without promoting them, use permission sets. They are additive bundles you can grant to a person, an API key, or an AI agent. See [Permissions](https://firefight.app/docs/gateway/permissions.md). --- Source: https://firefight.app/blog/introducing-firefight # Introducing FireFight, Open-source incident management, built for Slack ## The outage is only half the problem It's 5pm on a Friday. Half the team is already thinking about the weekend. Then checkout starts throwing 500s, and the support inbox fills up with people who can't pay. You find out because someone tags you in a thread that's already sixty messages deep. You scroll up and try to reconstruct it. Three conversations are tangled together in there. Someone is debugging, someone else is asking a question that got answered forty messages ago, and a third person has spun up another channel you can't find. Nobody has written down what's been tried or ruled out. You still don't know who's leading this, or whether it's yours. Then it goes quiet, which is somehow worse. Is it fixed, or did everyone assume someone else had it? Nobody says either way. It's Saturday morning before someone tags you again, asking for an update you don't have. Eventually the service comes back. But it's not over. There's a postmortem to write, customers to tell what happened, and lessons to turn into fixes. And every bit of it starts by reconstructing the story from scratch, out of six scrolled-past threads, with half of it already forgotten. That's the problem. Not the outage. The chaos around the outage. FireFight fixes that. And it's **open source**. ## What FireFight is FireFight is incident management that lives inside Slack. Declare an incident from anywhere, and the confusion turns into structure. A severity, a status, an owner, and a timeline that writes itself. When the incident closes, you're editing the postmortem, not writing it from scratch. No new tab. No separate tool your team forgets to open. It's live today, and you can self-host it. ## We made it open source. On purpose. FireFight is AGPL-3.0. The whole thing is on GitHub, and it's not a stripped-down version with the good parts held back. That matters even if you never run it yourself. Incident data is some of the most sensitive you have, and your security team can read exactly what touches it instead of taking our word for it. If your requirements change, you can take it and run it on your own infrastructure with your own model. [View the source on GitHub](https://github.com/FireFightLabs/firefight) Most teams would still rather not manage infrastructure for their incident tool, which is what Cloud is for. We handle the hosting and the upgrades. The difference is you're choosing us, not stuck with us. ## Run the whole incident in Slack You never leave Slack. There's nothing new to learn, because it's the tool you're already living in during an outage. Declare with `/firefight` and FireFight creates the incident and its channel for you. From there you can do all of this. - Set severity, status, and type inline, and keep the summary current - Assign a lead, or take it yourself from the pinned actions at the top of the channel - Pull in responders, and escalate when it's stuck. They get a DM with an Acknowledge button, plus a nudge if it goes unanswered - Say when the next update is due, and get reminded when it is - Link related incidents, or mark one a duplicate and merge it - Close it, or reopen it if it comes back FireFight runs the incident channel for you. It's created when you declare, the topic shows the current severity, status and lead, and it's archived once you close. Public incidents are also announced in a shared #incidents channel, and that post stays current as things change. Anyone who wants to keep an eye on things can follow along there without joining the incident channel itself. Private incidents skip the announcement and get a private channel instead, same command and same flow. ## The timeline writes itself You don't stop firefighting to log what happened. Everything that happens to the incident is already on the record. Every status and severity change with what it was before, who took the lead, who got escalated to and whether they acknowledged, the files people shared, the alerts that fired, and the moment it was resolved. Each entry is stamped with who did it and when. When something in the chat matters, pin it and it lands on the timeline too, or react to it to turn it into a tracked **action item**, a **follow-up**, or a **shoutout**. By the time the incident is over, the record is already there. Nobody has to reconstruct anything. ## AI that does the work, on your terms FireFight uses AI where it saves you time and nowhere it shouldn't. - **Catch me up** summarizes the incident for anyone who just joined, so you stop re-explaining - **Postmortems** are generated from the timeline you built, then you edit them - **Live summaries** keep the current state accurate without anyone maintaining it It's model-agnostic, so you're not locked to one vendor. ## See everything outside the channel Slack is where you run incidents. The dashboard is where you understand them. - A filterable incident list with stats and MTTR - A full timeline and audit trail of every state change - A service catalogue, custom fields, and configurable forms and roles ## Connect it to what you already run FireFight fits into your stack instead of asking you to rebuild around it. - A public REST API to create and manage incidents from your own tools - An alert endpoint your monitoring posts to, with rules that decide whether an alert opens an incident, joins one that's already open, or just notifies someone - Outbound webhooks with automatic retries Rules can work out who to pull in from your catalogue, so an alert about a service reaches the team that owns it. Repeat firings collapse into one alert instead of a wall of them, and a storm becomes one incident. You can test a rule against a sample before you trust it. ## Where this goes next We're building toward something bigger. A system that any responder, human or AI agent, acts through safely, with a real record of what happened and real limits on what can be done. That's the roadmap, not today's release. But everything above is the foundation it's built on, and all of it works right now. ## Build it with us FireFight is open source, and we're building it in public. The roadmap, the issues, the rough edges are all out in the open. Follow along, tell us what's missing, and help build it. [Join the Slack community](https://firefight.app/slack) ## Get started Sign in and declare your first incident in under a minute. [Sign in with Slack](https://app.firefight.app/) Or clone the repo and run it on your own infrastructure. Either way, the next outage doesn't have to be chaos. --- Source: https://firefight.app/changelog/2026-09-09-channel-archiving-delay # Choose when incident channels are archived (2026-09-09) **Channel archiving.** The Workspace settings page now lets you choose when Firefight archives an incident's Slack channel after the incident is resolved or cancelled. Pick anything from immediately to 30 days, or choose Never to keep every channel open. The default stays at one hour, and reopening an incident still unarchives its channel. --- Source: https://firefight.app/changelog/2026-08-28-dashboard-polish # Dashboard polish (2026-08-28) **Consistent tables.** Every table in the dashboard now uses the same layout, the sidebar shows the workspace you are in, and a row on the incidents list opens the incident wherever you click it. **Redesigned sign-in pages.** The sign-in, install and invite pages have been redesigned around a single form. **Reconnect Slack banner.** If Firefight loses access to your Slack workspace, for example because the app was uninstalled, a banner appears on every dashboard page. Admins get a Reconnect Slack button, and everyone else is told to ask an admin. The dashboard keeps working in the meantime. --- Source: https://firefight.app/changelog/2026-08-28-full-timeline-in-slack # The full timeline in Slack (2026-08-28) `/ff timeline` used to stop after the first few dozen events. It now pages through the whole history, with Older and Newer buttons and a caption showing where you are, such as "Showing events 51 to 95 of 140". A View full timeline button opens the same incident on the dashboard. --- Source: https://firefight.app/changelog/2026-08-28-routing-attribute-roles # Alert routing follows attribute roles (2026-08-28) Renaming a catalogue attribute used to be able to break paging. Routing now targets a role, not a name. In the type editor, mark an attribute as Members, Manager or Notification channel, and alert routing uses whichever attribute carries that role. Existing workspaces are already marked and route exactly as before. **Missing role warning.** If a routing rule depends on a role no attribute carries, a banner says so on Settings and on Alert Routing, and the route tester explains what is missing. --- Source: https://firefight.app/changelog/2026-08-28-postmortem-conflicts # Postmortem edits no longer overwrite each other (2026-08-28) A postmortem is often edited by a person and an agent at the same time. Saving one over the API or MCP now requires the version you edited. If someone changed it in the meantime, the save is refused and nothing is lost. **Dashboard editor.** When the dashboard detects that the document changed underneath it, autosave pauses and a banner offers to reload. **API change.** `version` is now required alongside the content when updating a postmortem. A stale version is answered with a conflict, so a client can fetch the latest and try again. --- Source: https://firefight.app/changelog/2026-08-28-configure-over-api-and-mcp # Configure the whole workspace over the API and MCP (2026-08-28) Statuses, severities, incident types, roles, alert sources, webhooks, API keys, agents, catalogue types, custom fields, runbooks, forms, routing rules, permission sets, grants and approval rules can all be created, changed and removed from the REST API and from MCP tools. **Postmortems.** Start a postmortem over the API or MCP, optionally with an AI draft, then read it, update it, and set its status. An agent that starts one is recorded as its author. **Routing dry runs.** Test a hypothetical alert against your routing rules over the API, the way the route tester does in the dashboard, without creating or notifying anything. **Member attributes.** A catalogue member attribute accepts an email or a member id. Someone Firefight does not recognise is refused with a message naming the value, rather than being silently added to the workspace. --- Source: https://firefight.app/changelog/2026-08-28-agents-first-class # Agents as first-class members (2026-08-28) The new Agents screen under Gateway is where you create one, issue and revoke its tokens, and disable or re-enable it. Agents act under their own name on the timeline, on action items, and as the author of a postmortem they wrote. **What agents can do during an incident.** An agent with the right grants can create, assign and complete action items, claim runbook steps, link related incidents, escalate, invite responders, and give a shoutout, the same things a person does from Slack. **Transcript access.** An agent can read what people actually said in the channel only if it holds a separately grantable transcript permission and the workspace has switched transcript access on. That switch is off by default. **Workspace settings.** A new Workspace settings screen holds the transcript switch, a retention window after which transcripts are purged (30 days by default), and the archive-channel setting. --- Source: https://firefight.app/changelog/2026-08-28-approval-rules # Approval rules you define (2026-08-28) The Ability Gateway has its own home. A new Gateway section in the sidebar holds Approvals, Activity and Permissions, with a count of pending approvals on the Approvals link. **Approval rules.** Under Permissions you now write approval rules. Each one says which abilities, risk levels and environments it covers, who approves (owners, admins, or named people and agents), where they are asked (the incident channel, a DM to each approver, or both), and whether a requester may approve their own request. Rules run in order and the first match wins. Until you write one, nothing waits. **Slack and the dashboard follow the same rules.** Grants and approval rules now govern Slack and the dashboard, not only the API. A Slack action caught by a rule is parked and completes on its own once approved. On the dashboard, a granted ability unlocks the matching controls, and a member without it sees the list without the buttons. **Request source.** Activity and Approvals show the source of every request, whether web, Slack, API or MCP. --- Source: https://firefight.app/changelog/2026-08-28-ai-timeline-notes # AI notes on the timeline (2026-08-28) When an incident resolves or is canceled, Firefight AI reads the conversation and adds those moments to the timeline as notes, each with the quote, who said it, and a link to the message, placed at the time it was said. **Where notes appear.** Notes appear in `/ff timeline`, in `/ff catchup`, and in the postmortem draft, so the write-up starts from what actually happened. Nothing is posted to Slack. **Dismissing notes.** Dismiss any note from the dashboard, and it collapses out of the way. Your own systems can subscribe to the new `milestone.noted` webhook event. --- Source: https://firefight.app/changelog/2026-08-28-run-incidents-from-dashboard # Run an incident from the dashboard (2026-08-28) You can now run an incident from its dashboard page. Resolve, cancel or reopen an incident, update its status or severity, assign the lead and every other role, escalate, invite responders, and give a shoutout, all without switching to Slack. **Same forms as Slack.** Every dialog renders the forms you configured, including custom fields, so a status update from the dashboard asks exactly what it asks in Slack. **Timeline entries link to what they mention.** Entries now show what they are about. A linked runbook opens the runbook, a related incident links to it, people appear with their avatars, action items appear as cards, and a pinned message shows the quoted text with a link back to Slack. When a runbook attaches automatically, the entry says which condition matched. --- Source: https://firefight.app/changelog/2026-08-28-github-native # GitHub, connected natively (2026-08-28) Connecting GitHub now takes you straight to the GitHub App install screen, where you pick the repositories Firefight may see. There is nothing to copy or paste, and no token to store. **Code tools for agents.** A connected agent can look up a pull request or a commit, fetch a slice of a file with the surrounding lines, search the code, and see who last changed a line and when. Each is a read-only capability you grant like any other. **Secret files are excluded.** Files that look like credentials, such as `.env` files and private keys, are refused outright and filtered from search results. **Automatic health checks.** Firefight checks the connection on its own, so the status on the integrations page reflects whether GitHub is reachable right now. --- Source: https://firefight.app/changelog/2026-08-10-ability-gateway # The Ability Gateway (2026-08-10) Give one engineer query access to the production database, and an AI agent the ability to file Linear issues, without either being able to touch anything else. The Ability Gateway is how you hand out exactly the access each person and agent needs, watch every use of it, and take it back the moment you want to. **Grants.** **Developer → Permissions** is where access lives. Grant a single ability, such as reading databases in PlanetScale, or a permission set that bundles abilities under a name like Database read-only. Scope any grant to specific environments, so one person reaches only Development while another reaches Production too. A grant can also carry an expiry date and end on its own. **Nothing inherited.** Service keys start with no access at all. An agent or script holds exactly the grants you gave it, never the reach of the person who created it. Revoking takes effect immediately. **Approvals.** Actions you mark as high risk do not run on a grant alone. They wait in **Developer → Approvals**, where an admin approves or denies them with full context of who asked and what they asked for. **The activity ledger.** Every attempt is recorded in **Developer → Activity** before it runs, allowed or denied. When something was refused, the ledger shows what was asked and why it did not happen. --- Source: https://firefight.app/changelog/2026-08-10-agents-can-act # AI agents can act, not just read (2026-08-10) Your AI agents can now do real work in Firefight, not just look things up. Ask an agent to update a runbook after an incident, fix who owns a service, or tighten an alert routing rule, and it makes the change itself, with the same permissions and approvals as everyone else. **Twelve write tools.** An agent with the right permission can maintain the catalogue (`upsert_catalog_entry`, `delete_catalog_entry`), manage alert routing (`upsert_routing_rule`, `delete_routing_rule`, `update_routing_config`), write runbooks (`upsert_runbook`), attach one to an incident (`attach_runbook`), shape what responders are asked (`upsert_custom_field`, `upsert_form_field`), hand out incident roles (`assign_incident_role`), and work the approvals queue (`approve_approval`, `deny_approval`). **Slugs, not IDs.** Everything is addressable by name. An agent can set which team owns a service, or gate a form field on the affected service being `checkout`, without ever looking up an internal ID. **OAuth for interactive clients.** Clients like Claude Code connect with one command and no token handling. The consent screen names the client asking for access and the workspace it would reach, and if you belong to several workspaces you pick one right there. Native desktop clients are supported too. **Governed like everyone else.** What an agent may change depends on the credential it connected with. OAuth connections act as you, service keys hold only explicit grants, and every call lands in **Developer → Activity**. --- Source: https://firefight.app/changelog/2026-08-10-integrations # Integrations, connected in one click (2026-08-10) Incidents get solved in the tools you already run, the code, the dashboards, the tickets, the database. Integrations connect those tools to Firefight, so responders and AI agents can read from them and act in them without leaving the incident. The catalog lives under **Configure → Integrations**. **What you can connect.** GitHub and GitLab for code, Linear for issues, New Relic, Datadog and Grafana for telemetry, Sentry for errors, Notion and Confluence for knowledge, and PlanetScale, Neon, Supabase and PostgreSQL for databases. Anything else that speaks MCP connects as a custom server with a URL. **One click for most.** Tools that host their own servers connect with **Continue with**, an approval on their own consent screen, and nothing to copy or paste. The rest take a URL and a token, stored encrypted and never shown again. **You choose what it can do.** Firefight discovers every capability the tool offers, and each enabled capability becomes a permission you can grant to people and agents. Capabilities that change something in the tool are marked **write**, and **Reads only** keeps a connection to the safe subset in one move. **A kill switch.** The toggle on a connected integration immediately removes its capabilities from everything, including agents connected over MCP, without losing your configuration. --- Source: https://firefight.app/changelog/2026-08-10-conditional-forms # Forms that ask only what applies (2026-08-10) Declaring an incident should take seconds. The dialogs responders fill in now adapt to the incident, so a routine declaration stays short and the important questions still get asked when the stakes are high. **Conditional fields.** Any field on a form can carry a condition, and it only appears when the incident matches. Conditions can read the incident type, the severity, and the values of other fields, including catalogue references like the affected service. A "Customer communication owner" field that only appears for your highest severities keeps low-stakes declarations fast while still capturing what matters when it counts. **Live in the dialog.** Conditions apply from the moment the dialog opens and re-evaluate as answers change, so responders only ever see the questions that apply to what they are declaring. **Leaner defaults.** New workspaces now start with forms that ask only what a new team actually needs. Fields that are always required, like severity and status, cannot be gated behind a condition, so a form can never hide what an incident cannot exist without. **For agents too.** `get_form` shows an agent everything a form holds, hidden fields included, and `upsert_form_field` changes visibility, requirements, and conditions with the same guardrails as the dashboard. --- Source: https://firefight.app/changelog/2026-08-10-every-role-assignable # Every role, assignable (2026-08-10) When something breaks, the first question is who is handling what. You can now assign every role you have defined, like Operations Lead or Communications Lead, during the incident itself, not just the Incident Lead, so anyone joining can see who owns each part of the response. **One dialog for the roster.** `/ff roles` opens a dialog with one picker per role, each showing whoever currently holds it. Change any of them and save, and Firefight posts a single message summarising what changed. `/ff lead` still works for handing over just the lead. **Visible where it matters.** Every assigned role appears in a Roles panel on the incident page, next to the person holding it. Role changes land on the incident timeline with the person named, so looking back you can see who took what and when. **Everywhere else.** Role changes reach your own systems as `role.assigned` and `role.unassigned` webhook events, and a connected agent can manage the roster with `assign_incident_role`, under the same rules as in Slack: one person per role, and the lead can be handed over but never left empty. --- Source: https://firefight.app/changelog/2026-08-10-runbooks-step-by-step # Runbooks, worked step by step (2026-08-10) A procedure only helps if it gets worked. When a runbook attaches to an incident, its steps are now a live checklist the team works through together, right in the channel. **Take a step.** The runbook message in the channel lists every step with an **I can take this** button. Press it and the step becomes an action item assigned to you, and the row turns into **Mark as done**. Complete it and the row is struck through with your name on it. The message updates in place, so a long procedure never floods the channel. **Hand steps out.** **View runbook** opens the full procedure with a person picker on each step. Assigning a step posts its action in the channel naming the assignee, with the controls to work it right there. **Progress you can see.** Each completed step posts a short line naming the runbook it came from, so progress through the procedure is visible without scrolling back. Only steps someone actually takes become actions, so the incident's action list reflects work done, not the whole procedure. **A sharper edge on conditions.** A runbook with no conditions no longer attaches to every incident. Attachment now only happens when conditions genuinely match. --- Source: https://firefight.app/changelog/2026-08-10-alerts # Alerts, routed into incidents (2026-08-10) Point your monitoring at Firefight and the alerts that matter become incidents, while the noise stays out of your channels. You decide what matters with routing rules, and Firefight handles the repeats, the flapping, and the storms. **Sources.** Create an alert source under **Settings → Alert Sources** and point your tool at its URL. The generic webhook works with Datadog, Grafana, Prometheus Alertmanager, or anything that can POST JSON, with a configurable mapping from the payload to normalized fields like `title`, `service`, and `severity`. Northflank has a dedicated integration. **Routing rules.** Rules run in order and the first match wins. Each pairs conditions on the alert's fields with one of four outcomes: create an incident, attach to an open incident, notify a channel or person without an incident, or drop. Alerts that match nothing are stored as unmatched and stay visible, so you can see what fell through and tighten your rules. **The catalogue fills in the blanks.** An alert carrying only `service: checkout` is enriched before routing. The owning team is merged in from the catalogue, and entry attributes become fields like `service.tier`, so a rule can route by metadata the alert itself never sent. **Built for noisy nights.** Repeat firings update the existing alert instead of creating new ones, a flap window catches alerts that resolve and immediately re-fire, and a grouping window attaches related alerts to the incident already open for the problem. One bad deploy that fires fifty alerts becomes one incident with fifty attached alerts, and one Slack digest message that updates in place. **Dry runs.** Test a hypothetical alert against your real rules from **Settings → Alert Routing** before pointing production at it, and see the per-condition trace of what matched. Connected agents get the same thing through `evaluate_routing`, which never creates or notifies. --- Source: https://firefight.app/changelog/2026-07-22-runbooks # Runbooks (2026-07-22) Your team's "here's what to do when this breaks" knowledge now lives in Firefight, and shows up exactly when it applies. **Structured procedures.** A runbook has a summary, a rich procedure written in a formatting editor and stored as Markdown, an optional external link, and an ordered list of steps. Manage them at **Settings → Runbooks**. **Automatic attachment.** Conditions decide which incidents a runbook applies to: incident type, severity, and custom fields, including catalog references like the affected service. Conditions are evaluated at declaration and re-evaluated on every update, so a runbook scoped to a service attaches the moment that field is set. **From checklist to actions.** Attached runbooks post into the incident channel with their steps and an **Add steps as actions** button. One press creates an action item per step, assignable and trackable like any other. **Built for AI agents.** Attached runbooks are part of the incident context Firefight's AI sees, and connected agents can browse them over MCP with two new read-only tools, `search_runbooks` and `get_runbook`. The REST API exposes the same data at `GET /api/v1/runbooks` under a new `runbooks` permission resource. **On the record.** New `runbook.attached` and `runbook.applied` timeline events, both available to outbound webhooks. --- Source: https://firefight.app/changelog/2026-07-10-incident-response-in-slack # Declare and run incidents in Slack (2026-07-10) Run the entire incident from Slack. Declare with `/firefight` and FireFight sets up the incident and its channel for you. From there you can: - Set severity and status inline - Assign a lead and pull in responders - Escalate when it's stuck, and acknowledge when someone picks it up - Close it, or reopen it if it comes back FireFight manages the incident channel through the whole lifecycle: created when you declare, kept current as things change, and archived on close. --- Source: https://firefight.app/changelog/2026-07-10-ai-postmortems-and-summaries # AI postmortems, catch-up and live summaries (2026-07-10) FireFight puts AI on the busywork around an incident: - **Postmortems** are drafted from the timeline, then you edit them - **Catch me up** summarizes the incident for anyone who just joined - **Live summaries** keep the current state accurate without anyone maintaining it The AI layer is model-agnostic, so you're not tied to a single provider. --- Source: https://firefight.app/changelog/2026-07-10-reactions-to-actions # Turn reactions into actions, follow-ups and shoutouts (2026-07-10) You don't stop firefighting to log what happened. React to a message and FireFight captures it: - Turn a message into a tracked **action item** - Capture a **follow-up** for after the incident is over - Give a teammate a **shoutout** in the moment they earned it The timeline builds itself out of how your team already communicates. --- Source: https://firefight.app/changelog/2026-07-10-incidents-dashboard # Incidents dashboard (2026-07-10) Slack is where you run incidents. The dashboard is where you understand them. - Filter incidents by severity, status and search - See stats including MTTR across your workspace - Open any incident for its full timeline and audit trail --- Source: https://firefight.app/changelog/2026-07-10-service-catalogue # Service catalogue, custom fields and forms (2026-07-10) Configure FireFight to match how your team works: - A **service catalogue** with typed attributes and relationships between entries - **Custom fields** on incidents - Configurable incident **forms, types and roles** --- Source: https://firefight.app/changelog/2026-07-10-public-api # Public REST API (2026-07-10) Bring incidents in from the rest of your stack. The REST API at `/api/v1` covers: - Creating and updating incidents, with a `source` so you know where each one came from - Reading severities, statuses and incident types - Reading and writing the catalogue Authentication is Bearer-token via API keys with per-resource permissions, and incident creation is idempotent, so retries never create duplicates. --- Source: https://firefight.app/changelog/2026-07-10-outbound-webhooks # Outbound webhooks (2026-07-10) Subscribe your own services to incident events. Deliveries are signed and timestamped, and failed deliveries are retried automatically, with delivery tracking so you can see exactly what was sent and when. --- Source: https://firefight.app/changelog/2026-07-10-incident-details-page # Incident details page (2026-07-10) Every incident has a dedicated page in the dashboard showing its timeline, its actions and follow-ups, and its postmortem, alongside the activity happening in Slack. --- Source: https://firefight.app/changelog/2026-07-10-timeline-attachments # File attachments in the timeline (2026-07-10) Files and screenshots shared in the incident channel are rendered inline in the incident timeline, so the visual context of what happened stays with the record. --- Source: https://firefight.app/changelog/2026-07-10-multi-workspace # Multi-workspace switcher (2026-07-10) If you belong to more than one workspace running FireFight, you can switch between them from a single account without signing out. --- Source: https://firefight.app/changelog/2026-07-10-open-source-self-host # Open source and self-hostable (2026-07-10) FireFight is open source under AGPL-3.0. Self-host it on your own infrastructure, keep your incident data inside your own perimeter, and point the AI layer at your own model. Because the codebase is open, your security team can read exactly what it does instead of taking our word for it.