CloudWarrior CloudWarrior, home

Run it yourself

Twenty-four questions.

Review recovery, deployments, monitoring, costs and access. Mark the answers you can support with evidence and leave the remaining questions open for investigation.

Work through the list with the person responsible for the environment. The answer count does not establish platform safety. Even one gap may need urgent attention; an unanswered question needs investigation.

This checklist supports an infrastructure discussion; it is not an audit or certification. Verify answers against documentation, configuration or tests rather than memory alone. To keep a copy, use the browser print dialog and save as PDF.

Recoverability

Can the platform come back without heroics?

  1. Can the last deploy be reversed by one person, without a meeting?

    Every rollback becomes an incident, and the fear of rolling back slows every release that follows.

  2. Has a restore from backup been performed in the last ninety days?

    An untested backup is a belief, not a control. The first real restore is the wrong time to discover the gap.

  3. Is the environment described in code that can rebuild it from nothing?

    Recovery time becomes however long it takes to remember, and the memory belongs to whoever is on holiday.

  4. Does anything in production exist only because someone clicked it once?

    That resource is invisible to review, absent from the plan, and the first thing to break silently.

Delivery

How does a change reach production?

  1. Can a new engineer ship a one-line change on their first day?

    Onboarding cost compounds. If the path is undocumented, every hire pays the same tax.

  2. Is the same artefact promoted through environments, rather than rebuilt per stage?

    The thing tested is not the thing shipped, so staging stops being evidence.

  3. Does the pipeline fail the build on a failing test, or is the gate advisory?

    An advisory gate is a comment. Coverage decays without anyone deciding to let it.

  4. Are deploys routine enough that nobody schedules them for a quiet evening?

    Batching changes to reduce risk raises it: bigger releases fail harder and are harder to attribute.

Visibility

Who finds out first when it breaks?

  1. Does an alert reach a human before a customer does?

    Support becomes the monitoring system, and the first signal arrives already angry.

  2. Can a request be followed across services without reading four dashboards?

    Diagnosis time grows with the number of services, which is the opposite of what the architecture promised.

  3. Does every alert have an owner and a runbook, or do some just fire?

    Unactionable alerts train people to ignore actionable ones.

  4. Is there a written record of the last three incidents and what changed after them?

    Without it the same incident is paid for repeatedly and each time it feels novel.

Cost

Does anyone know where the money goes?

  1. Can this month's bill be split by team, service or environment?

    Nobody can act on a single number. Attribution is what makes cost a decision instead of a complaint.

  2. Is there an alert on the invoice, not just on the infrastructure?

    Cost incidents are found at the end of the month, by which point they have already happened.

  3. Do non-production environments switch off when nobody is using them?

    Idle staging is the most common line item nobody has ever defended out loud.

  4. Has anything been right-sized in the last two quarters?

    Instances are provisioned for the worst day of the first month and never revisited.

Access

Who can do what, and who checked?

  1. Is there a standing human account with production write access?

    One phished credential becomes total compromise, and the audit trail cannot tell what happened.

  2. Are secrets held in a manager, or in environment variables and a pinned message?

    Rotation becomes impossible, so it never happens, so exposure is permanent.

  3. When someone leaves, is there one place that revokes their access?

    Access outlives employment, quietly, and nobody discovers it until an audit or a breach.

  4. Does the pipeline hold credentials that a person could also use directly?

    The blast radius of automation becomes the blast radius of every laptop.

Continuity

What happens when the person who knows leaves?

  1. Is there exactly one person who understands how the platform fits together?

    That is not a staffing risk, it is a single point of failure with a notice period.

  2. Do runbooks exist for the tasks that only happen quarterly?

    Rare tasks are the ones nobody remembers, performed under pressure, usually at the worst time.

  3. Are architectural decisions written down with the reason, not just the outcome?

    Every past decision gets relitigated, and the reasoning is reconstructed from guesses.

  4. Could the team operate for a month without any external help?

    Dependency on a supplier is fine until the supplier is unavailable and the dependency was undeclared.

What to do with your answers

  1. Review the impact of gaps

    For each no, identify the possible consequences and an owner. Missing recovery tests or access controls may matter even when there is only one no.

  2. Resolve the unknowns

    An unanswered question is not a satisfied requirement. Gather documentation, inspect configuration or plan a safe test with the system owner.

  3. Agree priorities

    Order the work by impact, urgency and dependencies. A full set of yes answers does not prove the absence of risk or replace an audit.

A consultation can walk through the noes against the running system rather than from memory, and agree which of them belong in a sprint.

What do you want to build or improve?

Tell us your goal and what is getting in the way. Two or three sentences are enough to start.

Tell us about your challenge

  1. Prepare an email
  2. Finish sending in your email app

Where the reply goes.

A product name or link is useful too.

What do you want to achieve, what is in the way, and is there a deadline? No passwords, keys or customer data.

Initial qualification 08:00–16:00, consultations with Patryk after 18:00, Europe/Warsaw. The date is agreed individually; sending this form does not book a consultation.

We use your details to handle this enquiry. You will not be added to a newsletter. Privacy.

Scripts are off, so this form cannot hand the request on. The same three answers work as plain mail: address, company and what needs to change. Write to pat@cloudwarrior.io.

Or write straight to pat@cloudwarrior.io.

Who answers, and when

Status: Patryk is currently contracted on a project.

  • Andrzej and Zosia run the initial qualification, 08:00–16:00 Europe/Warsaw. They ask about the company, the project, the problem, the outcome you want, the deadline and the budget, then pass the case to an engineer.
  • Technical consultations with Patryk are held after 18:00, Europe/Warsaw.

The time is agreed by email, case by case.

We will match the topic to the right specialist and work out whether you need advice, a focused delivery engagement or support for your team.

How we start

  • Andrzej or Zosia gathers the context and helps arrange the next step.
  • A consultant with relevant expertise reviews the challenge. If we are not the right partner, we will say so.
  • Before work begins, you receive a proposed scope, cost and timeline.
Contact
Directly with the CloudWarrior team
Next step
Advice or a delivery proposal matched to the challenge

Contact

Remote across the EU · Europe/Warsaw time