Skip to content
Kitchen Soap

Thoughts on systems safety, software operations, and sociotechnical systems.

Context and Operational Metrics

I really don’t think it can be overestimated how important context can be when it comes to troubleshooting or evaluating the health of an infrastructure. When starting to troubleshoot a complex problem, web ops 101 “best practices” usually start with asking at least these questions: When did this problem start? What changes, if any, (software, […]

Read more Context and Operational Metrics

Some Things We Did Today

Moving one of our eight photoserving farms from hardware Layer7 URL hash balancing (expensive, has limits) to L4 DSR balancing with CARP (cheap and simple) and figuring out how to juggle 18,000 requests/second while we do it. Built yet some more automated query analysis reporting (with some yummy MySQLProxy) Added yet another aggregated graph of […]

Read more Some Things We Did Today

Speaking at Web2.0 Expo 2009

Looks like I’m gonna talk about even more nerdy things at the Web2.0 Expo in April. You don’t have to wait for a recession to tighten up your operations. Squeezing more oomph out of your servers (or instances!) is always a good thing, and streamlining how you handle site issues is too. We’ll will talk […]

Read more Speaking at Web2.0 Expo 2009

Code Swarm for Config Management

Gil Raphaelli, one of the guys on our Flickr Ops team, put together a Code Swarm animation for the configuration/deployment management tool we use at Flickr to manage our infrastructure. Myles Grant did this for our bug reporting system as well. Check it out: Our automated config management system is called Gemstone, but conceptually you […]

Read more Code Swarm for Config Management