Meta tags:
description= sre content on DEV Community;
keywords= software development, engineering, sre;
Headings (most frequently used words):
the, and, to, sre, what, engineering, was, nothing, start, we, had, wrong, that, site, reliability, posts, dev, community, one, instance, bad, at, its, job, designed, notice, aws, field, manual, part, framework, sli, slo, sla, error, budget, latest, is, not, time, window, cold, never, practised, kubernetes, troubleshooting, check, before, you, restart, pod, could, without, dependency, filed, as, optional, backyard, endurance, os, designing, zero, loss, telemetry, ingestion, for, athletes, distributed, systems, node, js, consumer, capacity, webhook, delivery, rate, limits, dead, letter, decisions, ai, will, be, sometimes, then, go, media, pipelines, separating, lifecycle, validation, from, image, metadata, indexing, ten, minute, outage, took, six, hours, finish, edtech, password, recovery, auditing, global, logout, with, session, inventory, my, auto, remediation, bot, fixed, service, here, taught, me, about, real, your, uptime, monitor, keeps, crying, wolf, expand, migrate, contract, only, database, migration, pattern, small, team, needs, trending, guides, resources,
Text of the page (most frequently used words):
the (26), sre (26), follow (16), and (15), min (15), read (15), comment (15), comments (15), sep (15), dev (12), add (12), sergey (12), shinder (12), what (9), for (8), reliability (8), that (7), database (6), your (6), engineering (6), had (5), was (5), kubernetes (5), wrong (5), monitoring (5), nothing (5), devops (5), with (4), community (4), about (4), from (4), outage (4), posts (4), latest (4), reaction (4), sergeyshinder (4), start (4), create (3), software (3), postgres (3), one (3), why (3), without (3), agent (3), you (3), uptime (3), observability (3), monitor (3), time (3), systems (3), taught (3), node (3), troubleshooting (3), job (3), consumer (3), hive80 (3), lab (3), pingvera (3), com (3), lucas (3), ferreira (3), kenjitanaka6849 (3), architecture (3), media (3), carterhughes6853 (3), antonio (3), lopes (3), correia (3), rasmusberg6592 (3), run_as_daemon (3), opsforged (3), ivan (3), rossouw (3), enes (3), guler (3), menu (3), site (3), account (2), log (2), their (2), 2026 (2), post (2), teams (2), same (2), production (2), authentication (2), need (2), kafka (2), lag (2), ability (2), sort (2), top (2), relevant (2), sign (2), expand (2), migrate (2), contract (2), only (2), migration (2), pattern (2), small (2), team (2), needs (2), keeps (2), crying (2), wolf (2), auto (2), remediation (2), bot (2), fixed (2), service (2), here (2), real (2), edtech (2), password (2), recovery (2), auditing (2), global (2), logout (2), session (2), inventory (2), ten (2), minute (2), took (2), six (2), hours (2), finish (2), pipelines (2), separating (2), lifecycle (2), validation (2), image (2), metadata (2), indexing (2), will (2), sometimes (2), then (2), capacity (2), webhook (2), delivery (2), rate (2), limits (2), dead (2), letter (2), decisions (2), backyard (2), endurance (2), designing (2), zero (2), loss (2), telemetry (2), ingestion (2), athletes (2), distributed (2), could (2), dependency (2), filed (2), optional (2), check (2), before (2), restart (2), pod (2), cold (2), never (2), practised (2), not (2), window (2), aws (2), field (2), manual (2), part (2), framework (2), sli (2), slo (2), sla (2), error (2), budget (2), instance (2), bad (2), its (2), designed (2), notice (2), search (2), place, where, coders, share, stay, date, grow, careers, made, love, 2016, ruby, rails, built, powers, other, inclusive, communities, open, source, forem, terms, use, privacy, policy, code, conduct, mlh, shop, free, contact, showcase, organization, accounts, advertise, help, education, tracks, videos, challenges, home, space, discuss, keep, development, manage, career, push, credential, been, broken, days, commit, blocking, deleted, live, file, single, api, key, comparison, compatible, chat, integration, fintech, saas, apps, art, writing, good, mortem, archive, multiplier, eth_call, historical, block, upgrades, downtime, liveops, rollback, planning, when, game, event, goes, incident, postmortem, template, guide, has, access, problem, website, line, bash, script, ready, alerts, versu, mean, daemon, failed, doing, under, counts, outages, over, states, length, handoff, contracts, missing, piece, google, review, cheat, sheet, github, retries, microservices, circuit, breakers, how, them, future, next, years, look, like, effective, call, rotations, lessons, building, fair, schedules, cost, attribution, shared, infrastructure, worker, background, queue, retry, exhaustion, roadmap, day, plan, new, every, should, learn, little, rust, cron, systemd, timers, daemontools, understanding, evolution, linux, scheduling, mysql, connection, pooling, explained, concurrent, users, connections, replacing, exporter, modern, alternative, trending, guides, resources, wordpress, webdev, automation, security, llm, java, queues, webhooks, tutorial, incidentresponse, testing, cloud, loadbalancing, right, left, 157, older, hide, principles, practices, culture, close, powered, algolia, navigation, skip, content,
Text of the page (random words):
site reliability engineering dev community skip to content navigation menu search powered by algolia search log in create account dev community close site reliability engineering site reliability engineering principles practices and culture follow hide create post older sre posts 1 2 3 4 5 6 7 8 9 75 157 posts left menu sign in for the ability to sort posts by relevant latest or top right menu one instance was bad at its job and nothing was designed to notice sergey shinder sergey shinder sergey shinder follow sep 13 one instance was bad at its job and nothing was designed to notice sergeyshinder sre reliability loadbalancing comments add comment 2 min read aws sre field manual part 9 sre framework sli slo sla error budget engineering enes guler enes guler enes guler follow sep 13 aws sre field manual part 9 sre framework sli slo sla error budget engineering sre devops cloud monitoring comments add comment 4 min read latest n is not a time window ivan rossouw ivan rossouw ivan rossouw follow sep 13 latest n is not a time window observability devops testing sre comments 1 comment 4 min read the cold start we had never practised sergey shinder sergey shinder sergey shinder follow sep 11 the cold start we had never practised sergeyshinder sre reliability incidentresponse comments add comment 2 min read kubernetes troubleshooting what to check before you restart a pod opsforged hq opsforged hq opsforged hq follow sep 11 kubernetes troubleshooting what to check before you restart a pod devops kubernetes sre tutorial comments add comment 2 min read nothing could start without the dependency we had filed as optional sergey shinder sergey shinder sergey shinder follow sep 12 nothing could start without the dependency we had filed as optional sergeyshinder sre reliability architecture comments 1 comment 2 min read backyard endurance os designing zero loss telemetry ingestion for athletes and distributed systems run_as_daemon run_as_daemon run_as_daemon follow sep 12 backyard endurance os designing zero loss telemetry ingestion for athletes and distributed systems sre architecture observability postgres comments 1 comment 5 min read node js consumer capacity webhook delivery rate limits and dead letter decisions rasmusberg6592 rasmusberg6592 rasmusberg6592 follow sep 11 node js consumer capacity webhook delivery rate limits and dead letter decisions webhooks queues sre 1 reaction comments add comment 6 min read ai will be wrong sometimes what then antonio lopes correia antonio lopes correia antonio lopes correia follow sep 11 ai will be wrong sometimes what then ai java llm sre comments 1 comment 3 min read go media pipelines separating lifecycle validation from image metadata indexing carterhughes6853 carterhughes6853 carterhughes6853 follow sep 9 go media pipelines separating lifecycle validation from image metadata indexing go media sre comments add comment 6 min read the ten minute outage that took six hours to finish sergey shinder sergey shinder sergey shinder follow sep 9 the ten minute outage that took six hours to finish sergeyshinder sre reliability architecture 1 reaction comments add comment 2 min read edtech password recovery auditing global logout with session inventory kenjitanaka6849 kenjitanaka6849 kenjitanaka6849 follow sep 8 edtech password recovery auditing global logout with session inventory authentication security sre go comments add comment 6 min read my auto remediation bot fixed the wrong service here s what that taught me about real sre lucas ferreira lucas ferreira lucas ferreira follow sep 8 my auto remediation bot fixed the wrong service here s what that taught me about real sre automation devops kubernetes sre 1 reaction comments add comment 5 min read your uptime monitor keeps crying wolf pingvera com pingvera com pingvera com follow sep 8 your uptime monitor keeps crying wolf webdev monitoring wordpress sre 1 reaction comments add comment 3 min read expand migrate contract the only database migration pattern a small team needs hive80 lab hive80 lab hive80 lab follow sep 12 expand migrate contract the only database migration pattern a small team needs sre devops database postgres comments add comment 2 min read sign in for the ability to sort posts by relevant latest or top trending guides resources kafka consumer lag monitoring in 2026 replacing kafka lag exporter with a modern alternative mysql connection pooling explained do 50 concurrent users need 50 database connections cron vs systemd timers vs daemontools understanding the evolution of linux job scheduling s why every sre should learn a little rust the reliability roadmap a 90 day plan for new sre teams node js worker troubleshooting background queue retry exhaustion cost attribution in shared infrastructure effective on call rotations lessons from building fair schedules the future of sre what the next 5 years look like why your microservices need circuit breakers and how to add them what the github outage taught us about authentication retries google sre review cheat sheet agent handoff contracts the missing piece in production agent systems your outage monitor under counts outages and over states their length at the same time the daemon that failed by doing nothing monitoring versu i mean and observability website uptime monitoring from a 20 line bash script to production ready alerts your ai agent has the same database access you do that s the problem incident postmortem template guide for engineering teams liveops rollback planning what to do when a game event goes wrong kubernetes upgrades without downtime the archive multiplier why eth_call at a historical block the art of writing a good post mortem single api key comparison for compatible chat integration in fintech saas apps a push credential had been broken for 40 days the one commit it was blocking deleted a live file dev community a space to discuss and keep up software development and manage your software career home dev challenges dev videos dev education tracks dev help advertise on dev organization accounts dev showcase about contact free postgres database dev shop mlh code of conduct privacy policy terms of use built on forem the open source software that powers dev and other inclusive communities made with love and ruby on rails dev community 2016 2026 we re a place where coders share stay up to date and grow their careers log in create account
|