Meta tags:
description= In October, we experienced four incidents that resulted in degraded performance across GitHub services. This report also sheds light into an incident that impacted Codespaces in September.;
author= Jakub Oleksy;
Headings (most frequently used words):
github, availability, report, 2022, the, to, enterprise, in, insider, more, api, october, summary, newsletter, on, interested, bringing, your, organization, related, posts, explore, from, subscribe, product, platform, support, company, september, august, july, an, account, is, coming, all, customers, infinity, and, beyond, enabling, future, of, rest, with, versioning, exciting, new, features, powering, machine, learning, engineering, readme, project, actions, work, at,
Text of the page (most frequently used words):
the (86), #github (39), and (36), that (31), this (28), utc (25), our (22), incident (19), report (17), 2022 (16), codespaces (15), was (15), for (13), secret (13), not (13), were (13), which (12), october (12), from (11), more (11), availability (10), status (9), enterprise (9), new (9), also (9), have (8), impacted (8), change (8), engineering (7), are (7), all (7), experienced (7), resulted (7), services (7), impact (7), will (7), september (7), once (7), rotation (7), api (6), one (6), significant (6), minutes (6), blog (5), about (5), learn (5), still (5), after (5), into (5), service (5), error (5), hours (5), caused (5), regions (5), did (5), automated (5), lasting (5), webhook (5), back (5), activity (5), events (5), configuration (5), issues (5), time (4), their (4), future (4), with (4), what (4), accounts (4), your (4), month (4), jakub (4), oleksy (4), degraded (4), august (4), investigating (4), include (4), sheds (4), light (4), across (4), downstream (4), components (4), backlog (4), due (4), some (4), database (4), community (3), product (3), subscribe (3), coming (3), out (3), actions (3), machine (3), based (3), rest (3), user (3), trial (3), july (3), performance (3), incidents (3), multiple (3), resulting (3), pick (3), rotated (3), added (3), monitoring (3), term (3), updated (3), issue (3), component (3), missed (3), step (3), see (3), because (3), but (3), process (3), pools (3), jobs (3), port (3), forwarding (3), web (3), yellow (3), traffic (3), backend (3), healthy (3), receive (3), determined (3), data (3), worker (3), many (3), causing (3), webhooks (3), these (3), red (3), detected (3), cause (3), changed (3), organization (2), company (2), contact (2), support (2), developer (2), customer (2), stories (2), security (2), features (2), newsletter (2), insider (2), check (2), job (2), work (2), posts (2), team (2), exciting (2), learning (2), versioning (2), can (2), update (2), increased (2), account (2), customers (2), start (2), increase (2), free (2), two (2), next (2), follow (2), changes (2), working (2), identified (2), why (2), failed (2), additional (2), checklist (2), missing (2), steps (2), ensure (2), picked (2), longer (2), most (2), avoid (2), monitors (2), alerted (2), could (2), creation (2), earlier (2), system (2), restarted (2), west (2), europe (2), instance (2), quickly (2), saw (2), procedure (2), began (2), investigation (2), found (2), fail (2), reaching (2), state (2), mitigated (2), alerts (2), alert (2), improved (2), levels (2), disabled (2), normal (2), delivery (2), retry (2), any (2), source (2), find (2), large (2), however (2), had (2), unable (2), degradation (2), eliminate (2), older (2), type (2), deployment (2), region (2), global (2), during (2), schema (2), roll (2), rollback (2), metrics (2), green (2), values (2), production (2), errors (2), certain (2), value (2), alerting (2), systems (2), november (2), four (2), search (2), nov (2), privacy, terms, inc, linkedin, tiktok, twitch, youtube, facebook, twitter, shop, press, careers, training, forum, docs, desktop, electron, atom, partners, platform, resources, pricing, developers, covering, techniques, technical, guides, latest, innovations, current, openings, native, alongside, code, hosted, voices, readme, project, straight, explore, seth, juarez, discover, enhancements, empower, practitioners, powering, tim, rogers, introducing, calendar, keep, evolving, whilst, giving, integrators, smooth, migration, path, plenty, integrations, infinity, beyond, enabling, erika, melody, mileski, administrators, owners, responsibility, managing, keeping, secure, excited, introduce, soon, related, curious, other, plans, days, collaboration, per, expires, interested, bringing, acknowledges, com, june, get, best, directly, inbox, tags, please, real, updates, page, summary, then, identify, versions, use, easily, track, verify, rotations, address, short, verification, secrets, everywhere, automating, processes, human, several, later, separate, similarly, previous, failures, increasing, effect, immediate, used, exchange, token, cached, understanding, would, without, intervention, reality, only, heavily, drained, before, started, slowly, recover, accelerate, draining, queued, pool, size, helped, recovering, along, recovery, affected, performed, routine, received, internal, stating, client, statused, broad, observing, upon, few, expected, ran, feature, returning, moment, considered, investigated, rates, lack, overall, since, anomalies, cover, failure, mode, hour, brought, deliveries, got, workers, exist, fix, put, place, enabled, further, problems, encountered, recognize, must, resilient, spikes, load, make, improvements, learned, led, there, automation, creating, deleting, repositories, quick, succession, mitigation, order, give, solution, such, high, volume, rapid, series, create, delete, operations, triggered, influx, exceptions, needed, generate, payloads, been, deleted, attempting, tied, incoming, severe, queues, rely, them, delay, execution, severely, delayed, analyzing, dependency, repair, progress, entirely, verified, safe, practices, test, followed, individual, rollouts, discovered, cope, well, conflict, properly, tested, prior, rollout, additionally, version, does, gradual, exposure, worked, carefully, complexity, took, extended, period, complete, tracking, creations, rolled, propagated, variety, noticed, codespace, starting, trend, downward, deemed, enough, eventually, continued, following, mitigations, protect, against, testing, around, particular, area, fixed, dashboards, contained, inaccurate, pre, visible, deploy, help, prevent, initiated, steady, decrease, number, responses, apis, introduced, validation, required, present, correctly, set, default, every, scenario, population, null, cases, produced, when, pulling, records, projects, response, rate, went, within, traced, recently, deployed, recency, contributing, factors, provide, detailed, remediation, publish, first, wednesday, december, author, keyword, sales, policy, education, changelog, open, wayback, http, archive, org, 20221201185926, https, timestamps, capture, success, 2023, 2021, jan, dec, oct, 2025, captures,
Text of the page (random words):
github availability report october 2022 the github blog 21 captures 02 nov 2022 24 oct 2025 nov dec jan 01 2021 2022 2023 success fail about this capture timestamps the wayback machine http web archive org web 20221201185926 https github blog 2022 11 02 github availability report october 2022 blog engineering product security open source enterprise changelog community education company policy free trial contact sales search by keyword search engineering enterprise github availability report october 2022 in october we experienced four incidents that resulted in degraded performance across github services this report also sheds light into an incident that impacted codespaces in september author jakub oleksy november 2 2022 in october we experienced four incidents that resulted in significant impact and degraded state of availability to multiple github services this report also sheds light into an incident that impacted codespaces in september october 26 00 47 utc lasting 3 hours and 47 minutes our alerting systems detected an incident that impacted most codespaces customers due to the recency of this incident we are still investigating the contributing factors and will provide a more detailed update on cause and remediation in the november availability report which we will publish the first wednesday of december october 13 20 43 utc lasting 48 minutes on october 13 2022 at 20 43 utc our alerting systems detected an increase in the projects api error response rate due to the significant customer impact we went to status red for issues at 20 47 utc within 10 minutes of the alert we traced the cause to a recently deployed change this change introduced a database validation that required a certain value to be present however it did not correctly set a default value in every scenario this resulted in the population of null values in some cases which produced an error when pulling certain records from the database we initiated a roll back of the change at 21 08 utc at 21 13 utc we began to see a steady decrease in the number of error responses from our apis back to normal levels changed the status of issues to yellow at 21 24 utc and changed the status of issues to green at 21 31 utc once all metrics were healthy following this incident we have added mitigations to protect against missing values in the future and we have improved testing around this particular area we have also fixed our deployment dashboards which contained some inaccurate data for pre production errors this will ensure that errors are more visible during the deploy process to help us prevent these issues from reaching production october 12 23 27 utc lasting 3 hours and 31 minutes on october 12 2022 at 22 30 utc we rolled out a global configuration change for codespaces at 23 15 utc after the change had propagated to a variety of regions we noticed new codespace creation starting to trend downward and were alerted to issues from our monitors at 23 27 utc we deemed the impact significant enough to status codespaces yellow and eventually red based on continued degradation during the incident it was discovered that one of the older components of the backend system did not cope well with the configuration change causing a schema conflict this was not properly tested prior to the rollout additionally this component version does not support gradual exposure across regions so many regions were impacted at once once we detected the issue and determined the configuration change was the cause we worked to carefully roll back the large schema change due to the complexity of the rollback the procedure took an extended period of time once the rollback was complete and metrics tracking new codespaces creations were healthy we changed the status of the service back to green at 02 58 utc after analyzing this incident we determined we can eliminate our dependency on this older configuration type and have repair work in progress to eliminate this type of configuration from codespaces entirely we have also verified that all future changes to any component will follow safe deployment practices one test region followed by individual region rollouts to avoid global impact in the future october 5 06 30 utc lasting 31 minutes on october 5 2022 at 06 30 utc webhooks experienced a significant backlog of events caused by a high volume of automated user activity causing a rapid series of create and delete operations this activity triggered a large influx of webhook events however many of these events caused exceptions in our webhook delivery worker because data needed to generate their webhook payloads had been deleted from the database attempting to retry these failed jobs tied up our worker and it was unable to process new incoming events resulting in a severe backlog in our queues downstream services that rely on webhooks to receive their events were unable to receive them which resulted in service degradation we updated github actions to status red because the webhooks delay caused new job execution to be severely delayed investigation into the source of the automated activity led us to find that there was automation creating and deleting many repositories in quick succession as a mitigation we disabled the automated accounts that were causing this activity in order to give us time to find a longer term solution for such activity once the automated accounts were disabled it brought the webhook deliveries back to normal and the backlog got mitigated at 07 01 utc we also updated our webhook delivery workers to not retry any jobs for which it was determined that the data did not exist in the database once the fix was put in place the accounts were re enabled and no further problems were encountered with our worker we recognize that our services must be resilient to spikes in load and will make improvements based on what we ve learned in this incident september 28 03 53 utc lasting 1 hour and 16 minutes on september 27 2022 at 23 14 utc we performed a routine secret rotation procedure on codespaces on september 28 2022 at 03 21 utc we received an internal report stating that port forwarding was not working on the codespaces web client and began investigating at 03 53 utc we statused yellow due to the broad user impact we were observing upon investigation we found that we missed a step in the secret rotation checklist a few hours earlier which caused some downstream components to fail to pick up the new secret this resulted in some traffic not reaching backend services as expected at 04 29 utc we ran the missed rotation step after which we quickly saw the port forwarding feature returning to a healthy state at this moment we considered the incident to be mitigated we investigated why we did not receive automated alerts about this issue and found that our alerts were monitoring error rates but did not alert for lack of overall traffic to the port forwarding backend we have since improved our monitoring to include anomalies in traffic levels that cover this failure mode several hours later at 17 18 utc our monitors alerted us of an issue in a separate downstream component which was similarly caused by the previous missed secret rotation step we could see codespaces creation and start failures increasing in all regions the effect from the earlier secret rotation was not immediate because this secret is used in exchange for a token which is cached for up to 24 hours our understanding was that the system would pick up the new secret without intervention but in reality this secret was picked up only if the process was restarted at 18 27 utc we restarted the service in all regions and could see that the vm pools which were heavily drained before started to slowly recover to accelerate the draining of the backlog of queued jobs we increased the pool size at 18 45 utc this helped all but two pools in west europe which were still not recovering at 19 44 utc we identified an instance of the service in west europe that was not rotated along the rest we rotated that instance and quickly saw a recovery in the affected pools after the incident we identified why multiple downstream components failed to pick up the rotated secret we then added additional monitoring to identify which secret versions are in use across all components in the service to more easily track and verify secret rotations to address this short term we have updated our secret rotation checklist to include the missing steps and added additional verification steps to ensure the new secrets are picked up everywhere longer term we are automating most of our secret rotation processes to avoid human error in summary please follow our status page for real time updates on status changes to learn more about what we re working on check out the github engineering blog tags github availability report the github insider newsletter get the best of github once a month directly to your inbox subscribe more on github availability report github availability report september 2022 in september we experienced one incident that resulted in degraded performance across github services we also experienced one incident resulting in significant impact to codespaces we are still investigating that incident and will include it in next month s report this report also sheds light into an incident that impacted codespaces in august and an incident that impacted actions in august jakub oleksy github availability report august 2022 in august we experienced one incident resulting in significant impact to codespaces we re still investigating that incident and will include it in next month s report this report also sheds light into an incident that impacted codespaces in july jakub oleksy github availability report july 2022 in july we experienced one incident that resulted in degraded performance for codespaces this report also acknowledges two incidents that impacted multiple github com services in june jakub oleksy interested in bringing github enterprise to your organization start your free trial for 30 days and increase your team s collaboration 21 per user month after trial expires curious about other plans related posts enterprise an enterprise account is coming to all enterprise customers administrators or enterprise owners have the increased responsibility of managing their account and keeping it secure we are excited to introduce what is new with enterprise accounts and what is coming soon melody mileski erika xu engineering to infinity and beyond enabling the future of github s rest api with api versioning we re introducing calendar based versioning for our rest api so we can keep evolving our api whilst still giving integrators a smooth migration path and plenty of time to update their integrations tim rogers engineering exciting new github features powering machine learning discover the exciting enhancements in github that empower machine learning practitioners to do more seth juarez explore more from github engineering posts straight from the github engineering team learn more the readme project stories and voices from the developer community learn more github actions native ci cd alongside code hosted in github learn more work at github check out our current job openings learn more subscribe to the github insider a newsletter for developers covering techniques technical guides and the latest product innovations coming from github subscribe product features security enterprise customer stories pricing resources platform developer api partners atom electron github desktop support docs community forum training status contact company about blog careers press shop github on twitter github on facebook github on youtube github on twitch github on tiktok github on linkedin github s organization on github 2022 github inc terms privacy
|