Cloud reliability engineer reviewing CloudWatch alarm states and metric freshness dashboards at work

Your CloudWatch Alarm Is OK. Is the Metric Still Arriving?

Thu, Oct 8, 2026

Modern AWS deployment gates often rely on CloudWatch metric alarms to flag failures or stop rollouts. But what happens when a synthetic custom metric that should arrive every minute simply stops? CloudWatch can legitimately derive an OK, ALARM, or INSUFFICIENT_DATA state from missing telemetry, depending on the selected policy and prior state. None of those states, by itself, establishes when the last real sample arrived. The review therefore has to answer three separate questions: What did the alarm evaluate? When was the last real data bucket retrieved? Is the application actually healthy? This article proposes a bounded experiment in one authorized disposable AWS account and one explicit Region. It creates four action-free alarms, varies only TreatMissingData, stops and resumes a controlled publisher, and collects independent timestamps through GetMetricData. The output is an evidence ledger and one operational decision: ACCEPT, HOLD, REPAIR, or CHANGE-DESIGN.

The playbook is for AWS platform engineers, SREs, observability owners, and release-gate reviewers who already understand IAM, JSON, Python, and basic CloudWatch metrics. No AWS laboratory was executed while preparing this article. Documented behavior is presented as documented behavior; every cloud outcome in the proposed run remains an expectation until an authorized operator records it. The fixture never uses SetAlarmState, never changes production paging, and never claims that an alarm state proves application availability.

Define the claim an OK alarm can support

The experiment treats alarm state, metric freshness, and application health as different claims. A state of OK means only that the configured alarm evaluator currently classifies the metric according to its threshold, evaluation settings, retrieved datapoints, and missing-data policy. It does not mean a publisher is still alive. It does not mean the last bucket is recent. It does not mean users can reach an application. The fixture contains no application at all, so application availability is outside the evidence boundary.

The declared telemetry contract is deliberately simple: one synthetic publisher submits one value, either 0 or 1, in each distinct minute while it is running. Zero is nonbreaching; one is breaching. If publication stops, the test asks how each alarm interprets the silence and whether an independent query can still prove freshness. That contract differs from an intentionally sparse error metric, where silence may be expected. The canonical fixture remains continuous so that a missing minute is meaningful rather than ambiguous.

The operational labels are fixed before the run. ACCEPT means the required controls, silence phases, reconciliation, and fresh recovery were all observed. HOLD means the evidence is incomplete, ambiguous, or exceeded a local deadline. REPAIR assigns a defect in the publisher, query identity, or evidence pipeline without declaring the application healthy. CHANGE-DESIGN means the monitoring contract or policy is unsuitable and needs review. This freshness question is also different from the attribution question covered by AWS Network Synthetic Monitor documentation: a network health indicator can help localize an AWS-network impairment, while this experiment asks when a specific custom metric last produced a retrievable bucket.

Write the proposed gate claim in the manifest before creating the metric: “For this exact account, Region, namespace, metric name, RunId, unit, statistic, and completed query interval, the newest retrieved bucket is no older than the declared threshold, and the alarm state is consistent with its recorded policy and prior state.” That statement is intentionally narrower than “the service is healthy.” It also gives reviewers a concrete reason to reject attractive but irrelevant evidence, such as a dashboard from another Region, a similarly named metric without the RunId dimension, a successful application request that bypassed the publisher, or an alarm screenshot without the query interval. A deployment gate should fail closed on missing evidence, not expand its claim until the available artifacts appear sufficient.

Map the four missing-data policies without oversimplifying evaluation

The CloudWatch missing-data documentation defines four TreatMissingData strings: missing, ignore, breaching, and notBreaching. The setting matters only when the alarm does not have enough real datapoints to evaluate normally. With the default sliding evaluation window, CloudWatch can retrieve additional older datapoints. If it finds enough real points, it evaluates those points and does not use the missing-data treatment for that evaluation. This is why a recent gap does not automatically produce the all-missing outcome.

When every datapoint in the alarm evaluation range is missing, the conditional matrix is straightforward. It is a review aid, not a hand-written replacement for the CloudWatch evaluator:

Policy

Conditional all-missing state

What that state means

missing

INSUFFICIENT_DATA

CloudWatch cannot establish an evaluated threshold result from real datapoints in that range.

ignore

Prior state is retained

Silence preserves whatever state existed before the all-missing evaluation, including either OK or ALARM.

breaching

ALARM

Missing datapoints are treated as breaching for this alarm evaluation.

notBreaching

OK

Missing datapoints are treated as nonbreaching for this alarm evaluation.

AWS uses an illustrative five-point retrieval example in its documentation. That example must not be promoted into a universal promise about the alarm service's internal retrieval range. Sparse transitional cases have additional documented logic, and the service can continue reevaluating after the metric stream stops. The experiment therefore preserves every intermediate snapshot instead of demanding a fixed state exactly three or five minutes after publication ends.

A sparse transition is not equivalent to replacing every empty slot with the selected policy value. CloudWatch can combine real datapoints with treated missing datapoints, and the outcome depends on which real values remain available and whether enough of them satisfy the M-out-of-N rule. For example, the first capture after publication stops may still contain recent real zeroes or ones even though the latest completed minute is empty. That snapshot belongs in the timeline; it should not be discarded because it complicates the expected vector. The all-missing table becomes applicable only when the evidence supports that condition, and even then it explains the policy-derived state rather than the cause of the silence.

The DynamoDB namespace has a documented missing-data exception, but AWS-managed namespaces are excluded from this fixture. The alarms use a custom namespace so the four requested policy values remain directly comparable. The PutMetricAlarm API reference also documents that missing is the default when TreatMissingData is omitted. Here, every policy is supplied explicitly, and no alarm is updated midway through a phase.

Freeze one metric identity and four action-free alarms

Before any mutation, create a new run directory and record the AWS account ID, explicit Region, run UUID, source hash, aware UTC clock evidence, Python version, Boto3 version, botocore version, owner, local phase deadlines, and the 180-second demonstration freshness threshold. Persist the manifest first so an interrupted setup still reveals which names and identities may need cleanup. Refuse to reuse a directory or an existing alarm name.

The exact metric identity is Namespace=RefonteLab/AlarmFreshness, MetricName=SyntheticFailure, one dimension named RunId whose value is the run UUID, Unit=Count, and standard 60-second storage resolution. The four alarm names are derived from that UUID: refonte-freshness-<runid>-missing, refonte-freshness-<runid>-ignore, refonte-freshness-<runid>-breaching, and refonte-freshness-<runid>-notBreaching.

Use a purpose-built sandbox role rather than broad administrative access. The approved role needs only the actions required for this test: cloudwatch:PutMetricData, cloudwatch:PutMetricAlarm, cloudwatch:DescribeAlarms, cloudwatch:DescribeAlarmHistory, cloudwatch:GetMetricData, and cloudwatch:DeleteAlarms, plus caller-identity inspection through sts:GetCallerIdentity. Organization-specific permission boundaries, session controls, and network restrictions still apply. Do not silently attach AdministratorAccess to make the fixture easier.

The manifest is the ownership boundary for the run. It should include the intended alarm-name set before creation, the source-file hash, the command-line arguments, the person or automation identity responsible for cleanup, and the local deadlines for baseline, silence, and recovery. Write artifacts with collision-resistant names that include phase and observation UTC, and never overwrite an earlier raw response with a later compact summary. If initialization fails after only one or two alarms are created, the recorded name set still gives cleanup a bounded target. If the caller identity or Region changes between commands, stop immediately and classify the run as invalid rather than merging evidence from two environments.

Fix every variable except TreatMissingData

Create the CloudWatch client from one named, approved session and one explicit Region. Retrieve caller identity through the same session, write the manifest, check all four names for collisions, and create the alarms with identical settings except for TreatMissingData. The core pattern is:

from future import annotations

import json
import uuid
from datetime import datetime, timezone
from pathlib import Path

import boto3
import botocore

approved_profile = "approved-sandbox"
explicit_region = "us-west-2"
run_dir = Path("evidence") / str(uuid.uuid4())
run_dir.mkdir(parents=True, exist_ok=False)
run_id = run_dir.name

session = boto3.Session(profile_name=approved_profile)
cw = session.client("cloudwatch", region_name=explicit_region)
sts = session.client("sts", region_name=explicit_region)
identity = sts.get_caller_identity()

policies = ("missing", "ignore", "breaching", "notBreaching")
dimensions = [{"Name": "RunId", "Value": run_id}]
common = dict(
    Namespace="RefonteLab/AlarmFreshness",
    MetricName="SyntheticFailure",
    Dimensions=dimensions,
    Statistic="Maximum",
    Unit="Count",
    Period=60,
    EvaluationPeriods=3,
    DatapointsToAlarm=3,
    Threshold=0.5,
    ComparisonOperator="GreaterThanThreshold",
    ActionsEnabled=False,
    AlarmActions=[],
    OKActions=[],
    InsufficientDataActions=[],
)
names = {p: f"refonte-freshness-{run_id}-{p}" for p in policies}

manifest = {
    "run_id": run_id,
    "account_id": identity["Account"],
    "caller_arn": identity["Arn"],
    "region": explicit_region,
    "started_at_utc": datetime.now(timezone.utc).isoformat(),
    "freshness_threshold_seconds": 180,
    "python_boto3_version": boto3.__version__,
    "botocore_version": botocore.__version__,
    "alarm_names": names,
}
(run_dir / "manifest.json").write_text(
    json.dumps(manifest, indent=2), encoding="utf-8"
)

existing = cw.describe_alarms(AlarmNames=list(names.values()))
if existing.get("MetricAlarms") or existing.get("CompositeAlarms"):
    raise RuntimeError("Refuse to reuse an existing alarm")

for policy, name in names.items():
    cw.put_metric_alarm(
        AlarmName=name,
        TreatMissingData=policy,
        **common,
    )

Immediately export the resolved alarm configuration. Check the namespace, metric name, exact dimension list, statistic, unit, period, evaluation periods, datapoints to alarm, threshold, comparison operator, missing-data policy, ActionsEnabled=False, and all three empty action lists. Confirm that no conflicting evaluation-window mode is present. A mismatch is a setup failure, not an invitation to continue. The CloudWatch alarm guidance separates alarm state from actions and explains why a temporary SetAlarmState change cannot prove that metric evaluation generated a baseline; this fixture never calls it.

Build the publisher and an independent evidence collector

Use one small Python helper with init, emit, capture, history, and cleanup subcommands. Every command requires an explicit profile, Region, run directory, and phase. The run directory stores only manifest and evidence artifacts; it never stores credentials. emit accepts only integer 0 or 1 and a bounded count. capture uses a separately constructed CloudWatch client so the retrieval path does not depend on in-memory publisher state.

Keep the subcommands deliberately asymmetric. init may create names only after writing the manifest and collision check. emit may publish but may not delete or alter alarms. capture and history are read-only and must write the unmodified AWS response before producing a summary. cleanup may delete only names found in the manifest after it revalidates account and Region. This separation makes operator intent reviewable and reduces the chance that an evidence-gathering command mutates the system it is meant to inspect. Each command should write a small execution record containing argv, start and finish UTC, outcome, request IDs, and the tool versions that actually ran.

Publish real observations without backfilling gaps

For each publication, wait until approximately second five of a new minute when necessary, then submit the current aware UTC timestamp. Do not backfill a minute that was missed while the process was delayed or stopped. The PutMetricData API reference defines the timestamp, value, unit, dimensions, and storage-resolution fields. The submission pattern is:

from datetime import datetime, timezone

response = cw.put_metric_data(
    Namespace="RefonteLab/AlarmFreshness",
    MetricData=[{
        "MetricName": "SyntheticFailure",
        "Dimensions": [{"Name": "RunId", "Value": run_id}],
        "Timestamp": datetime.now(timezone.utc),
        "Value": value,
        "Unit": "Count",
        "StorageResolution": 60,
    }],
)

Log the phase, submitted timestamp, integer value, HTTP response status, and AWS request ID. If any call fails, stop the bounded publication loop and preserve the exception; do not invent a successful receipt and do not invisibly replay a request whose outcome is uncertain. When the publisher exits, append an explicit stopped_at event. The main process must not continue emitting while an operator believes a silence phase has begun.

Collect complete query responses before classifying absence

The collector queries the frozen identity with Maximum, Count, a 60-second period, ReturnData=true, and one query ID such as m1. For every capture, set StartTime to the manifest's minute-aligned run start and set EndTime to the current minute boundary, thereby excluding the open period. The GetMetricData API reference defines an inclusive start, an exclusive end, recent start-time rounding, timestamp ordering, and NextToken pagination. This explicit interval is selected by the experiment; it is not CloudWatch's hidden alarm evaluation range.

If EndTime is not later than StartTime, return NOT_YET_OBSERVABLE and wait for a completed interval. Do not issue an invalid query and do not relabel that condition as NO_DATA. Keep the same query bounds on every page. Preserve every raw page, query, status, and message. The MetricDataResult response model distinguishes Complete, PartialData, InternalError, and Forbidden. Intermediate PartialData is acceptable only while pagination is making progress toward a complete result. An unresolved partial response, repeated token, missing m1, unexpected message, internal error, or forbidden response produces UNKNOWN and HOLD.

Validate that timestamp and value arrays have equal lengths, pair them by index, sort chronologically, and deduplicate exact pairs. Reject conflicting values for the same timestamp. Never assume the last array element is newest without checking order. Complete describes retrieval for the requested interval at that time; a later, legitimately submitted historical datapoint could still change a future query. The collector therefore preserves both the query boundary and the observation time.

Establish OK using actual nonbreaching samples

Run init once, then publish zeroes in at least three distinct minutes. After the corresponding periods have completed, capture the metric and all four alarms. The baseline is valid only when the retrieved data contains real zero-valued positive controls and all four alarms are in OK. An initial alarm state or a manually forced state is not a baseline.

Make the sequence bounded and executable: run emit --value 0 --count 3, capture, and evaluate the criteria. If the criteria are not yet met, publish one new current-minute zero and capture again. Stop after the declared 12-minute local baseline budget. Preserve the timestamp of every iteration instead of treating a requested sleep duration as proof that CloudWatch evaluated. Re-export and compare alarm definitions after setup and before the next phase.

This pushed custom-metric experiment has different semantics from a Prometheus scrape. Refonte Learning's Prometheus and Grafana hands-on projects are useful for instrumentation practice, exporter design, and dashboards, but they are not authority for CloudWatch API behavior. Here, the positive control is the exact CloudWatch identity, the retrieved completed buckets, and the action-free alarm configuration. If the baseline does not converge inside the local budget, record HOLD rather than claiming an AWS failure.

The baseline packet should show more than four green state labels. It should contain at least three distinct completed bucket starts with numeric zeroes, a complete exact-identity query, the observation UTC used for freshness, and alarm snapshots captured after those buckets were available. Check that the retrieved values correspond to the intended publication receipts and that no value 1 appears in the baseline interval. If one alarm is already OK while the query cannot retrieve the zeroes, the state is not accepted as a positive control. If the query is complete but the alarms remain mixed, preserve that mismatch and continue only within the declared bound.

Stop publication and test silence from OK

Record the last successful zero submission and the publisher's stopped_at event, then verify that the publisher process has exited. Without publishing replacement values, capture metric pages and alarm snapshots once per minute for a local 15-minute silence budget. That budget is an experiment stop condition, not an AWS service-level commitment. Preserve every snapshot, including intermediate states that do not yet match the all-missing matrix.

Treat “publisher stopped” as an evidence claim of its own. Record the process exit status, the last attempted submission, the last successful request ID, and a process or job-control observation showing that no emitter remains active. A timer, second terminal, or scheduler can accidentally continue publication after the operator begins the silence phase. If a new bucket appears after stopped_at, do not call it an AWS anomaly until the extra publisher search is complete. Mark the phase contaminated, preserve the bucket and process evidence, and choose REPAIR or HOLD. The cleanest rerun uses a new RunId rather than trying to edit the contaminated timeline.

Keep the conditional expectation separate from a timed assertion

Once the observed timeline supports a sustained all-missing condition, the documented conditional expectation is:

  • missing -> INSUFFICIENT_DATA.

  • ignore -> retain the proven prior state, which is OK in this phase.

  • breaching -> ALARM.

  • notBreaching -> OK.

Do not demand that vector exactly three minutes, five minutes, or any other universal interval after stop. Older real datapoints and documented sparse-case logic can affect transitional evaluations. A quiet independent query over only the latest three periods also does not prove that the alarm service has no older datapoints in its own retrieval range. The experiment records what occurred and compares it with the conditional expectation only after the evidence supports that comparison.

The decisive counterexample is an OK alarm paired with stale independent telemetry. Preserve the query start and end, retrieval status, raw pages, newest bucket timestamp and value, calculated age, alarm configuration, and alarm snapshot together. That packet shows why “alarm is OK” and “metric is fresh” are different statements. It still says nothing about application health because the fixture has no application.

Repeat the silence test from a proven ALARM baseline

Resume publication with value 1 in distinct minutes until the collector retrieves real breaching buckets and all four alarms reach ALARM. With EvaluationPeriods=3 and DatapointsToAlarm=3, three breaching periods are necessary, but the experiment accepts the baseline only after the API evidence confirms it. Record the baseline timestamp, recheck all four definitions, stop the publisher again, and begin a second bounded silence capture.

For a sustained all-missing condition after the ALARM baseline, the conditional vector is INSUFFICIENT_DATA, retained ALARM, ALARM, and OK in policy order missing, ignore, breaching, and notBreaching. Compare the two ignore outcomes directly: it should retain OK after the first silence and ALARM after the second because the proven prior states differ. Preserve that prior-state evidence rather than attributing every ALARM observed during silence to a new breaching sample. If the vector remains unresolved at the local deadline, the decision is HOLD.

Measure freshness from retrieved bucket timestamps

Classify freshness independently of alarm state. With a complete exact-identity query, no returned buckets means NO_DATA. Otherwise, select the newest completed bucket's start timestamp and compute:

bucket_age_seconds = (observed_at_utc - bucket_start).total_seconds()

This is bucket age, not precise event age, ingestion latency, or publisher-process age. CloudWatch aggregates the sample into a one-minute period and returns the period-start timestamp. The capture excludes the open current period, so a healthy stream can still show a bucket that is more than a few seconds old. The manifest's 180-second threshold is a local demonstration tolerance, not a universal recommendation.

The calculation also depends on trustworthy clocks and an explicit completed-period rule. Preserve the collector host UTC reading, the unrounded observation UTC, the minute-aligned query end, and the returned bucket start. A future bucket start, a negative age, or unexplained clock disagreement must become UNKNOWN rather than being clamped to zero. Likewise, do not use the last PutMetricData client timestamp as a substitute for the retrieved bucket timestamp: the former proves what the client attempted to submit, while the latter is the evidence returned for the metric query. Keeping both lets a reviewer diagnose skew, delayed visibility, and identity mistakes.

Distinguish stale, missing, and unknown evidence

Use FRESH when age is between zero and 180 seconds inclusive, STALE when age is greater than 180 seconds, and UNKNOWN when retrieval is failed or incomplete, records conflict, the identity is wrong, the timestamp is in the future, or the observation clock cannot be trusted. Keep NO_DATA separate from a retrieved numeric zero. Zero is a valid nonbreaching observation; absence is not a value.

Alarm state

Freshness

Bounded interpretation

OK

FRESH

Supports a current, nonbreaching telemetry observation for the exact metric identity. It does not prove application health.

OK

STALE

The policy-derived state is OK, but the last retrieved bucket is too old for the declared freshness contract.

ALARM

FRESH

Supports a current breaching telemetry observation. Application impact still needs separate evidence.

ALARM

STALE

May reflect retained state, a missing-data policy, or earlier breaching data. Reconcile configuration and history.

Any state

UNKNOWN

Blocks a freshness claim because the independent retrieval evidence is not trustworthy.

Any state

NO_DATA

No bucket was returned for the complete exact-identity interval. Do not substitute a zero or call it stale.

This separation is the practical reason independent evidence matters. Refonte Learning's observability and API reliability guide discusses combining metrics, logs, and traces to understand a system, but the bounded claim here is narrower: the latest retrievable CloudWatch bucket must be timestamped and classified independently of the alarm state. An OK/STALE pairing is therefore evidence against a freshness claim, not evidence that the application is down.

Reconcile metric buckets, alarm snapshots, and history

Create one ledger row per run, phase, policy, and observation. Include observation UTC, metric namespace/name/dimensions, fixed query interval, page count, status and messages, newest bucket timestamp and value, bucket age, freshness verdict, alarm name, StateValue, StateTransitionedTimestamp, StateUpdatedTimestamp, StateReason, a pointer to preserved reason data, the proven prior baseline, the conditional expectation, the observed outcome, owner, and decision. The MetricAlarm response fields describe alarm state and evaluation state; they do not identify the newest metric sample. Preserve StateReasonData as an artifact without coding against an undocumented permanent JSON layout.

Do not compress away disagreement. Add fields for anomaly notes, source precedence, and the immutable artifact names that support each derived value. If the summary says STALE but the raw page contains a newer pair, the raw page wins and the classifier is REPAIR. If history and the current snapshot appear inconsistent, preserve both with their retrieval times instead of selecting the one that matches the expectation. A ledger row may therefore remain open even when most columns look normal. That is useful: it shows exactly which evidence is missing and who owns the next action, rather than turning an incomplete run into a generic pass or fail.

For each explicit alarm name and phase range, retrieve both state-update and configuration-update history through the DescribeAlarmHistory API. Follow every NextToken, store the original JSON, and derive compact tables only afterward. A missing transition does not prove that no evaluations occurred. Conversely, a StateTransitionedTimestamp or StateUpdatedTimestamp is not a metric timestamp and cannot replace the newest bucket from GetMetricData.

Keep the problem boundary visible. The Refonte Learning VPC Flow Logs capture-validation playbook treats SKIPDATA, eligibility, delivery, and query scope as a traffic-capture completeness problem. This article instead reconciles one custom metric stream with four alarm states. CloudWatch metric silence does not produce a VPC-style skip marker, so the ledger must preserve the absence, the query interval, and the policy interpretation rather than inventing a cause.

Diagnose mismatches without hiding the failed experiment

Use a fixed diagnostic order. First confirm account and Region. Then verify exact alarm names, namespace, metric name, the sole RunId dimension, unit, statistic, period, threshold, evaluation periods, datapoints to alarm, and action settings. Inspect publisher receipts and the stopped_at event. Check the completed query window, aware UTC clocks, pagination tokens, result status, all four configurations, and any unexpected publisher process. Only then compare the observation with documented transitional behavior.

Run one negative-control query with the same namespace and metric name but an intentionally wrong RunId. A successful complete retrieval should classify that exact identity as NO_DATA, not STALE, and must never substitute a neighboring series. If it returns a bucket, preserve the raw request and response and stop the run. The problem is identity or evidence isolation, not missing-data policy behavior.

Do not repair a mismatch by changing policies until a desired state appears. Export the discrepancy first. A local deadline is a stop condition, so an unresolved transition becomes HOLD rather than a claim that CloudWatch violated a provider guarantee. A fresh rerun, if approved, should use a new run UUID and retain the failed run as evidence.

Document source discrepancies explicitly as well. Living documentation can change, and a page may describe an example more narrowly than an operator first remembered. Record the access date, the relevant behavior in your own words, and the exact experiment assumption it supports. If two official pages appear to disagree in a way that affects the fixture, do not silently choose the convenient interpretation. Narrow the claim, capture both sources, and classify the affected acceptance criterion as HOLD until the owner resolves the scope. The purpose of the run is to make uncertainty visible, not to manufacture a clean demonstration.

Also separate publishing, discovery, retrieval, and alarm convergence. A successful PutMetricData request records that the service accepted the request; it does not prove when a completed bucket will be retrievable or when every alarm will transition. AWS notes that a newly created metric can take up to fifteen minutes to appear through ListMetrics, but this fixture bypasses discovery and queries a known identity. That listing delay is not an alarm convergence promise. The broader CloudWatch metrics concepts page explains identity, Region, dimensions, and period-start timestamps; the experiment still has to observe the requested data.

Restore fresh evidence and define rollback limits

Resume zero publication with current timestamps. Recovery requires new completed nonbreaching buckets, a local FRESH verdict, and eventual OK from all four alarms within the declared recovery budget. Do not backfill either silence interval. Fresh recovery establishes only the current telemetry condition and policy response; it does not erase or reinterpret the earlier gap.

If those criteria are observed, the run has a positive-control baseline, a silence phase from OK, a real-data ALARM baseline, a silence phase from ALARM, and fresh recovery. If any part is absent, keep the corresponding ledger row unresolved and apply HOLD or REPAIR as appropriate. Recovery is not permission to delete evidence or to declare the application healthy.

Define rollback limits before recovery starts. This lab has no alarm actions, production deployment, or customer traffic to roll back, so “rollback” means restoring the known publication phase and protecting the evidence boundary. Do not change the period, threshold, datapoints to alarm, statistic, dimensions, or missing-data policy during recovery. A configuration change would create a different experiment and could initially leave the state unchanged, making the timeline harder to interpret. If a configuration repair is necessary, close the current run with HOLD or REPAIR, export it, and initialize a new run with new names.

Preserve evidence before deleting the named lab alarms

Export final alarm configurations, alarm snapshots, every metric-data page, history pages, manifest, publisher receipts, stopped events, and the derived ledger. Stop the publisher. Confirm account, Region, run ID, and ownership before calling DeleteAlarms for exactly the four names recorded in the manifest. Verify their absence with an explicit-name DescribeAlarms call. Use the same recorded-name cleanup path after partial setup or interruption.

Custom metrics cannot be immediately deleted with a “delete metric” command. Cleanup therefore means ending publication and deleting only the four recorded lab alarms after evidence export. Do not delete unrelated alarms, mutate historical samples, or imply that a missing entry in a listing has erased the metric. The metric identity and custom-metric charges should be reviewed under the sandbox owner's normal cost controls.

Choose an alarm policy from the publication contract

Policy selection starts with the metric's publication contract, not with a universal ranking of the four strings. The action-free lab shows how state can diverge from freshness; production design must also define allowed lag, response ownership, and independent evidence.

Publication contract

Policy consideration

Independent control

Continuous once-per-minute telemetry

breaching may be appropriate when silence itself is an actionable telemetry failure; a separate freshness alarm may be clearer.

Query or heartbeat evidence must prove recency. Silence can also mean publisher or route failure, not application failure.

Intentionally sparse error signal

notBreaching may match a contract in which no event is normal.

Prove that the metric is intentionally sparse and monitor the emitter or route separately.

Unknown or disputed contract

missing exposes uncertainty as INSUFFICIENT_DATA rather than silently choosing good or bad.

Assign an owner to define expected cadence, allowed lag, and response.

Temporarily retained state

ignore preserves the prior state while data is missing. It can preserve either OK or ALARM and is not a permanent latch.

Record the proven prior state and establish how new real data resumes evaluation.

Keep a service-threshold alarm separate from a freshness or heartbeat check. A heartbeat that shares the same failing emitter, credentials, network route, or process has correlated blind spots. Moving it to a different path can reduce some shared failure modes, but a live heartbeat still does not independently prove application health. The owner must document exactly what each signal can and cannot establish.

No missing-data policy is universally safest. Refonte Learning's DevOps observability overview is useful background on metrics, logs, traces, Prometheus, Grafana, and operational monitoring, but AWS API behavior in this experiment comes from the AWS documentation and the captured evidence. Keep production paging, automated rollback, metric math, anomaly detection, composite alarms, and high-resolution designs outside this action-free lab review.

Apply acceptance criteria and assign operational ownership

Apply the decision labels to the evidence packet, not to a preferred narrative:

  • ACCEPT: Both real-data baselines were proven, both silence phases were captured, ignore demonstrated its retained-state difference, metric retrieval/configuration/history reconcile without material contradiction, and recovery produced fresh nonbreaching buckets with all four alarms back in OK. ACCEPT confirms this bounded telemetry experiment; it does not confirm application availability.

  • HOLD: A positive control is missing, a local deadline expires, retrieval remains partial or unknown, source behavior affecting the scope is unresolved, clocks or identities are ambiguous, or the evidence sources disagree. HOLD is the correct outcome when the claim cannot be supported.

  • REPAIR: The publisher used the wrong identity, failed to stop, backfilled gaps, or produced unreliable receipts; the collector used the wrong interval, lost pages, or mishandled status. Assign the repair to the responsible owner and repeat only with a new recorded run. REPAIR does not mean the application is healthy.

  • CHANGE-DESIGN: The policy does not match the declared cadence, the allowed lag is unsuitable, the freshness check is correlated with the failing path, or operators routinely read a policy-derived OK as proof of health. Revise the contract, signal separation, response, or alarm design through normal production review.

The proposed run sequence is therefore conditional, not historical reporting: publish zeroes and prove an OK baseline; stop and observe policy behavior from OK; publish ones and prove an ALARM baseline; stop and observe policy behavior from ALARM; then restore current zeroes and prove freshness. If the ledger supports every acceptance criterion, record ACCEPT. If any observation diverges or remains incomplete, retain the raw discrepancy and choose HOLD, REPAIR, or CHANGE-DESIGN.

The durable lesson is narrow and operational: alarm state alone is not proof of metric freshness. Pair every state used in a gate with an independent, exact-identity timestamp query and preserve the query boundary. Even then, the evidence speaks only to the metric stream and configured alarm. Application health requires its own controls.

About the Refonte Learning Cloud Engineering Program

Readers who want a broader cloud curriculum can review Refonte Learning's Cloud Engineering program, also called the Cloud Engineer Program on its page. The page describes a three-month schedule of approximately 12-14 hours per week. Its published scope includes AWS, Azure, and Google Cloud; infrastructure and networking; cloud security; virtualization and containers; Terraform and CloudFormation; serverless systems; automation and monitoring; and a capstone. It also states that mentorship and virtual internship opportunities are available, names MSc Charlotte Smith, and describes a Training Certificate and Certificate of Internship after successful completion.

The program page states degree prerequisites, while its FAQ says prior cloud experience is not required. Applicants who have already graduated should confirm the current eligibility wording directly with Refonte Learning. The page does not establish that this exact TreatMissingData/GetMetricData laboratory is included, nor does it verify Boto3 depth, cloud credits, sandbox allocation, mentor cadence, or an individual internship assignment. Review the curriculum as a broader learning path and confirm the availability of any specialized CloudWatch practice before enrolling. The program should not be presented as a guarantee of employment.