Tag: self hosted

  • Troubleshooting intermittent routing issues

    Quite some time ago, I started noticing random websites failing to load at random times. No real pattern to it. No obvious reason but also no lasting effects. Refresh a few minutes later and everything was back to normal.

    My browser? My network? Their network? All the new fancy cloud related technologies? Asynchronous microservices, serverless functions, NoSQL databases, ephemeral instances and everything else that happens when IT people decide to bring their commitment issues to work…

    There certainly were enough new technologies appearing. Too much scrumminess and agility up in the air. Way too much continuous integration, all the way to the back of the developer’s envelope. Teething trouble I thought…

    Looking at Firefox’s Network Monitor (Ctrl+Shift+E), the main culprits back then were fonts from fonts.googleapis.com and javascripts from one of the gstatic.com domains. Great! The undercover tracking “cookies” were now causing random sites to fail when unavailable. I blamed it on failing CDNs at the time and moved on as the problem gradually went away by itself.

    I may be connecting dots that were never there but some time later, google services would stop responding at times and I assumed the issue was back. Google search, google chat and youtube, all would randomly just stop working.

    The troubleshooting equations started with time and service as parameters. User, device and program were added very soon after. It started with the classic “works for me” responses and escalated to Cluedo style reports. “Google, on the laptop, with Edge is broken again.”

    Down the wrong rabbit hole

    I misinterpreted my trusty Firefox Network Monitor, which showed NS_BINDING_ABORTED errors. Those happen when the user is too impatient and clicks all over the place. Instead I assumed that the NS part meant DNS issues because of course, it’s always DNS.

    Was it Pi-hole? Was the Pi 2B finally dying? Or the 10 year old SD card? Is the latest version finally reaching the hardware limits? Am I even running the latest version?

    So many questions, one simple way to answer them all. Set up an instance in a Proxmox CT, sit back and wait. Did I mention that this process involved A LOT of waiting? The beauty of intermittent problems…

    Well, that didn’t fix it. I very briefly tried running without an ad blocker but that was not an acceptable option, especially not for a long waiting test. I tried AdGuard Home but that didn’t fix it either.

    I tried both Pi-hole and AdGuard with a combination of upstream DNS servers (ISP, google, quad9) as well as a local Unbound instance. Nothing made any difference. combinations returned the same results to the client so I went back to Pi-hole on the Pi 2B.

    At least I ended up with an upgrade to Trixie and valid repositories again. I also kept the Unbound instance, running on the same Pi, because why not add yet another potential point of failure in the chain.

    I was starting to suspect that this time it may actually not be DNS. I don’t know if that has ever happened before but I’m pre-emptively gonna claim first!

    It was around this time that I also started paying attention to the fact that NS_BINDING_ABORTED is not a DNS related issue.

    A more systematic approach

    One of the biggest pains in troubleshooting this, was that by the time I noticed, the time it took me to get annoyed and the time I tried to investigate, I only ever got a few minutes before the problem disappeared.

    I somehow realised that a simple ping was not a representative test to check connectivity to www.google.com. Ping (the program) selects one of the 8 IP addresses at random. So do all other programs, and if they don’t, DNS seems to return the results in random (maybe round-robin?) order.

    Did I mention that www.google.com has 8 A records? So does www.youtube.com. Some sort of redundancy I guess and a natural, external load balancer…

    dig www.google.com
    ...
    ;; ANSWER SECTION:
    www.google.com. 706 IN A 142.251.150.119
    www.google.com. 706 IN A 142.251.155.119
    www.google.com. 706 IN A 142.251.156.119
    www.google.com. 706 IN A 142.251.157.119
    www.google.com. 706 IN A 142.251.153.119
    www.google.com. 706 IN A 142.251.154.119
    www.google.com. 706 IN A 142.251.151.119
    www.google.com. 706 IN A 142.251.152.119

    I soon was able to confirm that whenever I was having trouble with a service, at least one of the IP addresses was not reachable. No ICMP responses, no traceroute, no response to an http request nothing.

    I needed a way to check this without having to very quickly type ping commands and it would certainly help to keep track of when it happens, how often and how long it lasts.

    Gatus

    Gatus to the rescue! I don’t know much about their hosted service. I needed to monitor google from my location so I went with a local docker instance and boy, has it delivered! Here is a gist of the config I used to simply ping the 8 IP addresses every minute.

    Is it still a plug if I praise someone else’s service? Who cares. I’m happy with the few features that I have used so far and I will most certainly be looking at it further. That feature list looks impressive!

    Any conclusions?

    What have I found out so far? Well, I’m not going (completely) mad. The red bars do exist but that’s as far as the insight goes for now.

    • They last a VERY consistent 15 minutes
    • They don’t seem to be random. Sometimes it’s all fine for a couple of days, sometimes they happen multiple times a day.
    • The outages by IP, seem to be (loosely) connected. Quiet days are quiet across the board. On busy days the outages seem to happen in sequence.
    Gatus dashboard from 22 September 2026
    Gatus dashboard from 22 September 2026

    If it was any other website or service, I would argue that their service rollout process is to blame, or their network update process, or their failover process or something along these lines.

    We are talking about google services however. That’s the part that does not compute. It’s not that they don’t have youngsters that very enthusiastically push the wrong button. It’s that if a technician has beans in the evening, the news reports “excessive methane levels in the data centre” the next day. It could be my ISP but then it would be the national news instead…

    That leaves 2 options. Either the issue is indeed on the side of my ISP but only affects a small subset of users, my network “neighbours” so to speak, or the issue is in my own network.

    More data is required and more may soon be coming. I’ve asked the question on my ISP’s community forum and people have already offered to monitor from their own connections. There will probably be a 2nd part to this with further findings.

    Alternate theory

    There is of course a slight chance that they intentionally bury all the google outage reports from the search results. That means that yet another one of my conspiracy theories has been confirmed and this blog will never be popular now.

    Anyone getting to find and read this entry constitutes and actual paradox and I need to stop writing and go find my tinfoil hat at once. So long and thanks for all the fish!

  • BI on a shoestring – an ode to materialised views

    Some background info

    As I may or may not explain on some other post, some other day, I run my home lab on hardware that’s well past the “obsolete” stage and marching full speed towards becoming “vintage”.

    Said lab, hosts all sorts of little services, private and public, including this here, marvellous website as well as a docker instance of BrickSync, a great little piece of software that keeps our BrickOwl and BrickLink stores in sync.

    On a routine search for my lost disk space, I (re)discovered that BrickSync spits out a snapshot of the store inventory in XML format every time an update was triggered. Either by an order or by me. It also saves a copy of every order in XML format.

    All in all, I discovered over 20 GB worth of XML files that contain a nearly full history of my store. A short python time later and I had a POC that could walk through all the files and dump all the data. Now what?

    The failed attempt

    A little voice in the back of my head had been shouting Elastic, which I hadn’t used before. A chat with a friend confirmed the idea. A ton of little files, ingested and turned into useful, searchable information. Exactly the task Elastic was made for.

    So, off I went to set up a stack on my poor little home lab. That’s when I was reminded that Elastic is meant to run on proper server hardware. I kept adding memory, trying to get it to do something useful, and I eventually got to that point with 4 GB of ram.

    Yes, it run and yes, it looked like it was doing something but not much and not fast. Compared to everything else on that server, its footprint looked like Bigfoot was stepped on by a dinosaur, and the spatter was causing OOM kills among the rest of the services.

    Back to the drawing board

    Then I thought my data is already very structured. Do I really get much benefit from using Elastic? How about pushing the data in some database, MariaDB being the prime candidate, and then treating it like the time series it is, maybe using Prometheus and Grafana.

    While researching the subject, the Postgres name kept popping up and I decided to give it a shot. I have a couple of instances running but never developed for it. A quick setup and a quick update to the python scripts and data was loading in the background while I continued researching the Prometheus – Grafana setup.

    As it turns out, Prometheus is itself a database. A Time Series Database to be precise and as it turns out, Postgres also has a time series extension, TimescaleDB. The things you find out if you RTFM first…

    By this point, I already had nearly 50 million records imported and I had already discovered that Grafana can actually pick up data directly from Postgres. I decided to skip the completely unfamiliar to me, Prometheus step, and leave that lesson for another day.

    My thought process in roughly this order:

    • This is SQL. I know this!
    • The Grafana query builder is a bit of a pain. I’ll create views in Postgres.
    • These views are kinda slow. Since I don’t have real time data to worry about, I could cache the view results into a table…
    • Hey look! Materialised views!

    Where were those when I was young? That’s similar to the SQL Server Indexed Views, right? Is that where the idea of cached views came from? This brain could do with some defragmenting and re-indexing…

    What did we achieve?

    What a fantastic thing those materialised views! Aggregate 50 million rows in 2, 3, 4 different ways, base your queries on those and get sub-second responses!

    Refreshing these views takes a few minutes on this beast of a system, but new data is only available a few times a day so daily updates are more than enough. Those will probably be scheduled to run during the night.

    The current setup consists for one Postgres and one Grafana instance, each allowed 512MB of ram. I can get the Grafana container killed if I open a bunch of pages back to back but I have not yet seen the Postgres container die.

    What’s next?

    The TimescaleDB extension might still come in handy. Incremental updates to the materialised views, time buckets and alleged space savings with columnar storage compression are definitely worth checking out.

    For the front end, something lighter than Grafana might be worth exploring. At the moment it feels like I’m using a sledge hammer to present data straight out of views. Maybe something based on Apache ECharts would be enough. A project for when we need to recover server resources.

    Why?

    So what does it do? Well it does exactly what it was meant to do. It provides very clear and valuable insights into the inner workings of an AFOL’s mind. It also highlights other mysteries of the lego world.

    How, for example, when black, white and grey parts only make up a quarter of the store stock, they account for over half of the sales. Or how when over half of the sales are of black, white and grey parts, all the things that people build and share are so colourful…

    Part Count by Color
    Part Count by Color over the last 5 years. X-axis could do with some more formatting
    Parts sales by color

    It does more and it can do even more. My next challenge is to determine what other data can be extracted and visualised, and more importantly, find out in what way is it of any value.

    In any case, finally getting to play with Postgres in a non-superficial way is a win in itself. It seems to have a very slightly larger idle footprint than MariaDB and it has been a very well behaved tenant on this limited server so far.

    Materialised views were a very nice, unexpected little surprise. Apparently Postgres has a full meme worth of such features waiting to be explored. Can’t wait!