GitHub Outage Tracker: Is GitHub Cooked?

(isgithubcooked.com)

90 points | by toomanyrichies 2 hours ago

14 comments

  • kashnote 1 hour ago
    I think we need to have a little more sympathy for GitHub. You could justify the jabs when we could all blame any outage on the migration to Azure, but then they shared numbers around the scale they're dealing with now that everyone is constantly building and pushing with AI.

    I think it's commendable that they're not limiting access to the site or (intentionally) throttling newcomers. Yes, they need to get this figured out, but a little sympathy goes a long way. I personally wish them the best and hope their on-call people can go back to getting normal amounts of sleep soon.

    • gogobio 49 minutes ago
      Sympathy? It's a Microsoft company that is being ran with a consistency of a startup in early seed rounds. Their downtime is abhorrent and unacceptable as far as enterprise goes. Their engineers look like absolute amateurs allowing for such low class work it results in their customers experiencing industry leading downtime.
    • herpdyderp 1 hour ago
      > I think it's commendable that they're not limiting access to the site or (intentionally) throttling newcomers.

      I don't think this is commendable at all. I give GitHub a lot of money and I'm tired of it being wasted with downtime.

      • jjice 57 minutes ago
        It's a shame that the GitHub org that we use at my job that we pay a lot of money for gets affected the same way my personal nonsense does.

        I don't know the architecture or any of that, but I feel like there could be (and it's not like they would've really known this until the last year or two with the massive spike) separate infrastructure for paid users/orgs vs free the same way they make the distinction with enterprise.

        I get the massive load changes that they are under over the last two years, but why does a bunch of vibe coded slop take down the same resources that my company pays for every single month and has for years? I imagine properly splitting that out would be an absolute headache and not worthwhile for them vs stabilizing the rest of the service, but damn it sucks when I get blocked at work because GH is down.

        • tempest_ 5 minutes ago
          Our GitlabCE instance has been sitting in the racks for nearly two years with almost 100% uptime running on 10 year old xeons that have long since paid for themselves.

          Of course it is not free of all management but for our use case it is working.

          There are hiccups with the CI runners from time time but nothing major and we have another machine in another rack that serves as a backup which can be brought up ~< 20 minutes.

          I know companies have long since tossed their expertise for hosting their own stuff in favour of SaaS but at some point its hard to beat the up time of a single machine.

    • mronetwo 39 minutes ago
      We have a business relationship so no we shouldn’t have any sympathy. They sell a service and they’re failing to provide it.
      • amelius 27 minutes ago
        Why? If we have sympathy for Apple then we can have sympathy for Microsoft ...
        • pimeys 14 minutes ago
          Wait, I don't sympathize Apple at all... Or any other American corporation.
        • mirashii 17 minutes ago
          You're the only mention of Apple in this thread, and I don't see what they have to do with it.
    • padjo 55 minutes ago
      I don't typically have sympathy for businesses that fail to deliver a service as advertised.

      I can have sympathy for the humans caught in the crossfire but only managing one nine of availability on a commercial service is not acceptable.

    • niltecedu 1 minute ago
      idk man, we pay stupid amounts of money to microsoft, we are an enterprise customer, expecting better availablity compared to my laptop isnt really a high bar.
    • dpz 33 minutes ago
      Maybe for a free account.

      But we pay enterprise license and GitHub is a big dependency in our software flow.

      If this continues to be a problem as an enterprise product they need to do something. Otherwise theyre are going to to start losing business

    • JDups 37 minutes ago
      Constant building with AI is something that they (Microsoft) promote and are heavily invested in.
    • theideaofcoffee 3 minutes ago
      Give me a break. Sympathy? For microsoft? That might have flown when github was like seven people, but they have nearly unlimited resources to make it better. They're just choosing not to. Let's talk contracts and money before we pull the sympathy card.

      I used to be on-call in a high-traffic environment where single customers pushed more bits than entire nations. I chose the role. I didn't want people's sympathy, if anything, I wanted them to complain to management.

      If it gets too bad they can quit. Maybe that would be for the best, just wear the thing down until it outright fails and no one wants to touch it.

  • JeremyHerrman 52 minutes ago
    > "GitHub has had 1125 incidents since February 2016, implying a monthly incident rate of 24"

    1125 incidents / 126 months ≈ 8.9 incidents per month, not 24

    still terrible, but why such an obvious error in the first sentence...

    • 6LLvveMx2koXfwn 17 minutes ago
      Not sure whether it has been updated since your comment, but the sentence now reads:

          GitHub has had 1125 incidents since March 2016. Over the last 3 months, they've averaged 24 incidents per month
      
      edit: although they also have 1.2 days of downtime (in a day) for their 'worst days' of downtime table, which suggests some auto number crunching is not working as expected.
      • gen220 5 minutes ago
        Yes I tweaked it! The number and copy were mismatched and are no longer!

        That worst day is likely an overlapping incidents accounting issue; I tried to account for overlapping incidents in another view but probably failed to port it over there.

        Should be fixed soon!

    • stevage 6 minutes ago
      > GitHub has had 1125 incidents since March 2016. Over the last 3 months, they've averaged 24 incidents per month (↓ 5% vs prev 3mo).

      Looks like they fixed it already

    • 4petesake 19 minutes ago
      Prob used Copilot to write the excel formula...
  • nightpool 43 minutes ago
    Getting rid of Actions and Copilot and other secondary services almost halves Github's incident rate: https://i.imgur.com/XPcMIFr.png

    I'm a big fan of Github Actions and I think people are often a little too harsh on it, but it's clear that it's sad that it's come at such a high cost to the platform's stability

  • bushbaba 1 hour ago
    could have been a page with a static 'Yes' and a significant portion of time it'd be accurate.
  • 404mm 1 hour ago
    If backend GitHub services are anything like GHES then I’m surprised it even managed to scale this much.
  • djieidj283 10 minutes ago
    It’s ironic that the SCM that 2026 software engineers landed on, is one with a complicated distributed usage model and is slow/down because of a centralised service
  • GrumpySciGuy 18 minutes ago
    Yes, but I did not need to look at the tracker to know that.....
  • rvz 54 minutes ago
    They were cooked the moment they got acquired by Microsoft.

    This is why I foresaw that centralizing everything to GitHub was just generally a bad idea 6 years ago. [0]

    Now that there is no CEO of GitHub, there is no point to GitHub improving.

    [0] https://news.ycombinator.com/item?id=22867803

  • fenio 1 hour ago
  • kevmo 1 hour ago
    An important thing to consider is how much of their uptime without incidents is not the normal working hours. Their incident-free uptime on 9-5 EST, Mon-Fri, is probably like 60%.
    • thombles 38 minutes ago
      As a daily GitHub user in Australia I still haven’t figured out why everyone’s complaining about uptime. :)
    • CoastalCoder 56 minutes ago
      > Their incident-free uptime on 9-5 EST, Mon-Fri, is probably like 60%.

      And it may be even worse in EDT, which is currently in effect!

    • perfectstorm 1 hour ago
      what's normal working hours? very US centric comment IMO. Europe, India, China, Latam etc. don't fall into your 9-5 EST normal working hour bucket.
      • padjo 50 minutes ago
        A service like GH will still show a daily usage pattern, often with a peak somewhere around 16:00 UTC when most of the US and Europe are at work.
      • boredatoms 26 minutes ago
        Im fairly certain that SWEs only exist on the US west coast
  • arlattimore 1 hour ago
    I literally chuckled out loud when I saw the headline :D
  • johnea 2 hours ago
    Oh, I thought it said "Github Outrage Tracker".

    I was ready to click...

    • HPsquared 1 hour ago
      That one would be like iscaliforniaonfire.com
  • jakub_g 1 hour ago
    Possibly inspired by:

    https://red-squares.cian.lol/

    • gen220 1 hour ago
      Hey! isgithubcooked.com is my site; I made this site in February '26, I think

      The contribution graph as an outage calendar idea is a commonly recurring one :). I definitely saw it somewhere else as a static asset first before I made this site.

    • ChrisArchitect 1 hour ago