← All episodes
EPISODE 140November 18, 2019 · 40:00

Kubernetes Best Practices

Read the transcript

Brendan: [00:00:00] It’s a powerful tool, but it’s also kind of a footgun.

Bridget: It’s time for Arrested DevOps, the podcast where we help you achieve understanding, develop good practices, and operate your team and organization for maximum DevOps awesomeness. I’m Bridget Kromhout, and we’ll introduce our guests after a word from our sponsors. The worst thing about the Arrested DevOps podcast is when it ends. You’re left wondering what to do next. What are you going to listen to on your commute home? How do you occupy your time when walking the dog? What are you going to listen to during the quarterly all-hands meeting? But fear not, dear listener, there is a solution. You need to subscribe to Software Defined Talk right now. It’s a weekly podcast that recaps all the news in cloud computing, DevOps, and enterprise software. The hosts, Cote, Matt Ray, and Brandon Wichard, will keep you up to date on all things cloud while offering tips on how to optimize your Costco haul, and how to PowerPoint. It’s a fun, free-flowing conversation that will keep you entertained and informed. What are you waiting for? Subscribe to the podcast today by visiting softwaredefinedtalk.com or by searching for Software Defined Talk in your favorite podcast app.

[00:01:18] Okay, it is super exciting to be chatting with all 4 authors of the forthcoming Kubernetes Best Practices book from O’Reilly Media. I’ll have them introduce themselves in, say, byline order. Who are you and what makes you want to write books about Kubernetes? Let’s start with Brendan, if we haven’t lost Brendan.

Brendan: I think that I’ve written a few books. I guess it’s a legacy of once upon a time being a professor. I really like to teach people. I guess maybe it’s a byproduct of the fact that I’m excited about people using and empowering them with technology. If you don’t teach them how to use it, that doesn’t usually happen. I write because I like to teach.

Bridget: That’s awesome. I love it. Okay. All right. Eddie.

Eddie: Hey, I’m Eddie Villalba. I’ve been at Microsoft for 10 years, actually last week, and just understanding that being able to really understand what customers are looking for, what people are looking for, and my knowledge in the product has really driven me to spread that wealth, that knowledge, to others. Dave and I both have been in the field for quite some time, and seeing what can go wrong, and epically wrong, and how we can possibly get that out there for even small startups that don’t have the luxury of getting help from companies like Microsoft to really look at this space of Kubernetes and cloud-native and get the knowledge that the big folks learn really quickly, and also from bad experiences and from good experiences. That’s why I love to get this written.

Bridget: [00:03:01] All right. Dave?

Dave: Yeah, my name is Dave Strebel. I help customers daily be successful with Kubernetes. I won’t lie, I never had aspirations to really write a book, but when the opportunity presented itself, things I do like to do is help break down complex technologies to make them more simpler, to help people really understand those technologies. So I was really glad that I did help write a book. So it really changed my mind, you know, how to really help people with technology.

Bridget: All right. And Lachie?

Lachlan: Hello, my name is Lachlan Evenson. I wanted to write the book specifically because I wanted a chance to give back to the community. I have learned so much from the community over the years and people, you know, I remember in the early days approaching Brendan in like late 2014, early 2015, and he stood in the hallway and answered all my questions. And that was kind of a really, you know, it’s the community steps up and answers these questions. So one is I wanted a chance to give back. The other one was I wanted an opportunity to write a book that I wanted. So I’ve been through the journey of Kubernetes and it growing up in the ecosystem. So I would have loved in 2015 to have this kind of almanac of all the things that you should do when looking at Kubernetes and kind of shortcutting all the decision points you might have to make when building and operating Kubernetes. I think this book provides that kind of level of, you know, get to the problems you need to answer really quickly. That was my excitement with having the opportunity to write the book.

Bridget: [00:04:45] Nice! For any of our listeners, if you acquire a time machine, instead of bringing a sports almanac to the past, you should bring this book. And give it to Lockie in 2015. All right, action item for all listeners, who all, of course, have to go get the book and read the book, and then get it and put it in your time machine. Okay, so I think for a lot of our listeners, they may be using Kubernetes, which is awesome. Some of them are like, that is all I hear about, but why? What even is this Kubernetes? Where did it come from? Brendan, you’re probably responsible for something. Why don’t you give us the 5-second elevator pitch? What’s going on with this Kubernetes thing, and why do we need best practices for it today?

Brendan: Well, I mean, I think that it has moved on from being something that people heard about somewhere to something that everybody is looking to implement. But I think that our experience building these managed services is that While people are convinced by the value that the tech can bring, they sometimes struggle in figuring out how to accomplish the particular task that they want to do. And, you know, as Dave said, we work with a bunch of people hands-on, but that’s not super scalable. And so, you know, I think that Writing down the best practices that we’ve seen helps sort of uplevel the knowledge in the entire community. Because it is important to do it right. We’ve seen lots of people sort of shoot themselves in their foot. It’s a powerful tool, but it’s also kind of a foot gun. So, we want to make sure people do it right. That’s why we write these things down. I think in contrast to other books that other people have written at times that sort of explain Kubernetes in general, I think what’s great about this book is it’s really focused around specific topics. We don’t expect you necessarily to read it cover to cover, but rather to dip in on a topic when you’re working on, say, machine learning, or, you know, you’re working on another aspect of the project and you want to just read a quick summary of, like, what should I do? You know, what should I be thinking about as I approach doing machine learning on Kubernetes, or what should I think about as I approach setting up a Kubernetes cluster for a bunch of developers. So it’s kind of maybe more of a series of short essays rather than a whole put-together book. I think that’s great. There was a need for that kind of consulting for people.

Lachlan: [00:07:24] Right.

Bridget: So people shouldn’t expect narrative flow or narrative structure necessarily, but this is something where they can go directly to what interests them. I’m curious, and maybe this is something that I know Lachie and I were talking about a little bit, what makes now the right time to do a best practices book?

Lachlan: For me, it was now there’s more adoption of Kubernetes out there in the ecosystem, and the ecosystem has become more complex as a result of everybody using it. There are a bunch of tools. Kubernetes itself has quite a sprawling variety of different APIs. So when it comes to how do I solve policy with Kubernetes, for example, it’s nice to have that— here’s where to look, here’s where to start, here’s how the community is approaching something like policy or a topic like policy— and give you all the pieces that constitute that specific topic. So I think we’ve covered in the book all those kinds of touchpoints as you go on your journey from Kubernetes, from deploying your first service to doing rolling upgrades, to looking at policy, to looking at governance. To looking at security, we’ve kind of covered in detail in each of the chapters these pieces. And I think that’s what the community needs right now. What are the aspects I need to worry about when using Kubernetes throughout my journey? So whether it’s your first day using Kubernetes or you’re in a large enterprise using Kubernetes and it’s like, I have this new thing, as Brendan said, it should be that reference you have on your desk and you can get value out of it at any point. But I think Short-circuiting the time it takes to make a reasonable decision about a specific topic is why Best Practices is out there now. Also, we’ve got some air miles on Kubernetes in the real world. Dave, Eddie, myself, Brendan, we’re out helping customers, and we’re seeing a wide variety of questions that come in. We’ve used that to guide how we wrote the book to give people that reference material as, here are the problems I’m going to need to solve. Reasonable ways to solve them. I just think that maturity in the ecosystem has led to this being the right time to write the book.

Eddie: [00:09:33] I want to echo Lockie. The maturity is there as well in the sense that there are a lot of organizations, the community at large, they’re already down the path of these projects and they really need— they don’t want the step-by-step walkthrough that a lot of— it’s all over the place, right? You see it on the internet everywhere. Hey, if I’m looking at this and I’m going to my project manager or going into a Scrum meeting, these are the few things we need to make sure that we look at while we’re focusing on a specific topic. There hasn’t really been a place yet for that, and I think this book really hits those points. It’s just a great little desk-side book to say, hey, I need to look at networking. All right, pop it open and get to that little spot right there, or policy, as Lakki was saying. Definitely the velocity of both the community and also the velocity at which Kubernetes is changing. I think we did a really good job, as well, to make sure that we covered topics based on what we know about the project and what we know about how it’s moving in the future, and try not to make it a snapshot in time of this version or that version of Kubernetes, but best practices that should follow through based upon the mission and the focus of what Kubernetes is going through. Being part of the community and being part of the contributors helps us keep that mindset as we’re going.

Dave: [00:10:52] Yeah, and to kind of Lachie’s point on, you know, the ecosystem really growing and becoming a lot more complex, I think there’s an understanding that users need to have to focus on those basic core concepts within Kubernetes. A lot of times, they try to skip over those and go on to kind of over-engineer their environments, layer on more complex technologies, But, it was really focused on understanding those core Kubernetes concepts that you need to learn and understand these things before layering on other technologies.

Lachlan: Just in hearing what everybody else said, I really like the mix of this philosophy, because people are signing on to, hey, we’ve gone to DevOps. We’ve gone cloud-native. This is the continuum journey. The book is a great mix of philosophy, so why you would want to do this, the problems it’s aiming to solve. And kind of tactical how you would solve them in Kubernetes. So there’s that mix that makes it a good read. So it’s not just how to do policy on Kubernetes. It’s, you know, for example, it’s why would I want policy and what am I actually trying to solve as part of this in that continuum of cloud-native ecosystem and moving your workloads to cloud-native?

Bridget: [00:12:10] Okay. So I feel like we’ve talked about your policy chapter a couple of times, and I do want to go into detail about a few of your favorite chapters, or at least some of the chapters, because you all wrote several chapters in this book. And we can’t, of course, give our listeners a thorough overview of every single topic. I have a PDF in front of me. It’s 258 pages. So let’s talk about your policy chapter specifically. Like, what stands out for you there? What should you kind of tell people that might make them say, oh, That’s for me.

Lachlan: Well, for starters, it’s Chapter 11 and everybody loves hearing Chapter 11 for anything. So that’s a great start. But I think in the ecosystem, as a lot more enterprises are coming on and looking at Kubernetes as a platform to move all their workloads, they have this outstanding question of how do I deal with policy and governance? And what this really is about is how do I make sure that the workloads that I deploy are conformant to some policy that we’ve set up, whether you’re in a heavily regulated environment or you just want to understand and how things are configured. That’s such a new concept in the ecosystem because obviously policy isn’t the first thing you address when building something like Kubernetes. But now we just see a real uptick in people asking about how do I deal with policy. It was fun to write that chapter from my perspective to give people an overview of where the ecosystem is at, where the tooling is at, and the things that you might be able to achieve with policy. That was my favorite chapter to write, although it was fun to write them all. But I think this answers a really, really salient question in the community right now, which is, how do I actually get control of my workloads and make sure that they are compliant? So I hope everybody appreciates the context that’s given in that chapter.

Bridget: [00:14:09] And I think that’s fantastic too, because if you think about it, If there is anything that enterprises that have actual customers and money care about, is they care about making sure that everything that they’re trying to control does get controlled in the expected way. It’s nice to see the Kubernetes community actually focusing on that. I know there’s the open-source project Gatekeeper in that space. I think that’s one of the examples that you go through in the chapter. If I’m not incorrect.

Lachlan: Yeah, that’s correct. And the great thing about something like Gatekeeper is it’s built on top of OPA, which is an open policy agent in the cloud-native ecosystem. So we have a generic policy controller or policy engine that people can use, and things like Gatekeeper make a Kubernetes-native implementation of OPA. So there’s a lot of ways policy can be expressed, but I think, you know, the high-level philosophy of how you might want to achieve policy on Kubernetes is illustrated, and then there’s kind of the tactical, how could I actually do this today using some open-source tooling in the ecosystem?

Bridget: [00:15:16] All right, and so other chapters that people were interested in highlighting. Dave, tell us about your favorite.

Dave: My favorite was resource management. It doesn’t sound really exciting at all, but it’s something that I see users struggle with a lot. With users I work with, and it has a lot of impacts on other things besides just running the actual workload, but it also has a lot of impact on things like scaling within Kubernetes. I really like that chapter because I think I learned a lot myself from it, things I didn’t know around technologies that were in Kubernetes that you should really think about more when deploying your workloads. That’s kind of why I I was excited about that chapter, even though it doesn’t sound exciting, but it really is, and it’s a core thing that you have to really understand before running workloads.

Bridget: [00:16:19] I mean, it actually sounds sort of exciting from the perspective of, if you’re on call for production, possibly resource overruns are what’s going to page you. People love the idea of having that be well controlled so that they don’t end up with, gosh, Wait, we had to save memory for the control plane. Interesting. I feel like there’s some— Kubernetes has some guardrails in there for that, right? Or is that kind of, set it how you like?

Dave: Yeah, there are definitely guardrails, and the book really kind of hits on those best practices around setting up things like requests and limits that have a huge impact on how you run workloads and how your workloads are going to behave. When you do run out of capacity.

Bridget: Wait, you’re saying that the cloud is not infinite? Interesting.

Lachlan: I wish I had this chapter, you know, back in 2015, because I would say that most of the outages I was paid for in the early days— I think most people trip across on resource management in the early days as that cluster gets to 80, 90, 100%, and it’s how do you deterministically understand what’s going to happen when things start to run hot is really great context for you to be able to put the guide rails in place very early rather than having that 3:00 AM page, which, you know, the cluster is going into cascading failure. It’s something that I’ve personally seen happen, and resource management would have been a great chapter for me to get started back in my early days of Kubernetes.

Bridget: [00:17:52] All right, Eddie, tell us about the chapter that you found. The most interesting to write?

Eddie: I found it not only interesting, but I think it’s probably the most challenging because I had to fit a lot of stuff into a concise format. Mine was chapter 9, which covers networking, network security, and the new magic buzzword, service meshes. I think the reason I like it a lot is because it’s the foundation. Without getting this right, without getting networking correct, The architecture begins to fail right away, especially when we start looking at hybrid platforms, when we start looking at cloud-based platforms and integrating it with very complex enterprise systems and complex WAN systems. Little, tiny things trip us up, right? DNS, the move from kube-dns or SkyDNS to CoreDNS and how that’s configured, that blew people’s minds the first day. Like, oh, this worked before, and now it’s not working. Little tiny things. Then, the other part, especially in the enterprise space, is we saw Kubernetes had this huge, just, oh, we’re playing with it, and now we’ve got to see it in production. In some cases, like the customer I’m working with, a lot of the divisions just decided to put it in and not even tell anybody, and then come back later, and security’s like, what is going on here?

Bridget: [00:19:16] What did you do?

Eddie: And that’s like, all my controls are gone? What do you mean? And why is this not being seen by our central security stack? And how do I get those? And that has all become— it’s kind of, in some enterprises, it’s come kind of full steam ahead in saying, we need to kind of treat this as if it’s just another node on our network. But when it comes to the Kubernetes space, there are new things and new paradigms that we have to understand, right? North-south traffic is now handled potentially inside of Kubernetes, not by a device. And, you know, I talked specifically around network policy agents and integrating with your CNI and some best practices around choosing those toolings. Then, once you have that foundation, everything looks hunky-dory, now, immediately, the next conversation is, well, someone told me I need a service mesh for my—

Bridget: You need a service mesh because it’s a floor wax and a dessert topping, right?

Eddie: Right, exactly. What’s the office supply company that had that little easy button? I don’t want to have to rewrite my app, but I want observability, I want security, and I want policy. I just want to press this little easy button. What they’re finding out, unfortunately, is it’s not so easy. Thousands of little buttons that you have to press in the right combination.

Bridget: [00:20:34] Is it a hard button?

Eddie: Reality is, it’s not that hard when you start to— Again, because service mesh is so new, and I did not want to make this a point-in-time type consideration. I try talking more about generalities of— everybody agrees that service meshes should do certain things correctly, right? If we decipher that and break those down into what those things are, and deciding why you need those things and what the priority is for your application stack, and then picking your vendor from there. Then I talk a little bit about the SMI spec, which kind of helps level that, and it’s that idea of, let’s get a common API against those common things that all service meshes should do. Then, as service mesh creators start writing modules, they say, hey, if I meet this, I know that I’ll have this available to everybody at this level, and then I can add my value adds from there. I kind of COVID that, and I want to make sure people are aware of of what to look for when they’re trying to decide something as critical as a service mesh for their technologies.

Bridget: [00:21:45] All right. Brandon, I know you don’t want to pick a favorite child, but if you want to give us some of your thoughts now that you’ve completed this book and the parts of it that you guided and led, what stands out for you? What should people definitely flip to that page and read?

Brendan: Well, I mean, the first chapter is sort of an intro. How do I even just lay out a service? So I think, you know, some people may come at the book having already done a bunch of the basics, but like if you come at it and you want to, you know, you’ve learned all these Kubernetes objects, but then you’re like, wait, okay, how do I actually put it all together? That chapter is a great one just to make sure people start on the same right foot. But I actually like the one that follows that probably the best, which is about how you set up developer and developer workflows on the cluster. Because I think that, you know, sometimes we adopt this tech and the operators love it, or it’s great for continuous delivery, but we’ve made the developers’ lives miserable. But I think there’s a lot that you could do to actually really use Kubernetes to not just make the operations and the running of the application better, but actually make developing applications easier as well. So, that chapter talks about, you know, how you could partition the cluster with namespaces and how you can, you know, take a look at the ways in which sort of the different lifecycle moments in using a cluster. There’s the act where you’ve just hired a new developer and you say, okay, you know, we’re using Kubernetes and let’s onboard you, let’s get you set up with the environment that you can use. There’s the aspect of, like, making sure that people can accidentally step on each other’s toes with RBAC and things like that. There is providing sort of cluster-level services for people so that they don’t have to learn about logging. Logging is just there. They don’t have to learn about monitoring. Monitoring is just there. And then things around sort of testing and debugging that are critical flows that change a little bit, right? You’re not just doing your development necessarily locally on your machine. But you’re using this cloud-based resource or you’re using this cluster-based resource, testing and debugging are different. And it’s important that we make sure that that’s easy for people too. I think especially with testing, testing, because if it’s not easy, people will do less of it, and then you ship buggier software, right? So it’s really a critical component of success is actually getting the environment set up before you ship the software or the place where you build the software I think it can be sort of an underlooked or an underappreciated part of the thing. We think a lot about production, but what goes into production comes from the place where we develop and the place where we test. That’s a chapter that I like a lot. Every chapter is great.

Bridget: [00:24:32] That’s awesome. That’s a really good point, actually, because, of course, my mentality is I focus on production, but if people who are creating the production experiences don’t have good onboarding and don’t have good reproducible environments, then how are they going to produce good production experiences? I like that, the idea of making Kubernetes workflows for developers, whether it’s shared clusters or instantiation at will or whatever it is that they’re doing, making that more reproducible, making that easier, is really valuable. I remember reading that section when I was originally reading through the book and thinking, Oh, there’s a lot of good ideas here. Maybe the hard question is, how do you get organizations to use them? Those of you who are out there dealing with people in the field, do you mostly see people doing the stuff that you’re recommending here?

Brendan: [00:25:34] Well, actually, I think we wrote the book, I wrote the book because I think people want to. I don’t think that they’re like, no, we really want to provide a horrible developer experience. Or, you know, we want to slow down our developers, or we want to do a bad job with service mesh. I think there’s a lot of desire, but there’s just not— the recipes just aren’t there, necessarily.

Eddie: I agree 100% with Brendan on that. I mean, the customer I’m focused on right now is going through that journey, and it’s twofold. It’s like, hey, we get to reset some of this bad stuff that we did before because of institutional processes that were put in place long before 98% of the developers that are working there now are there. They get to do a quick reset and say, now we’re going to this new cloud-native world, and how do we have this reproducible process? They even created their own little new division called— basically, they’re looking at patterns and practices, but they’re basically called Advocated Patterns Division. That whole job is just to say, How do we put patterns together for our developers, our ops folks, to say, our new world is all Kubernetes, it’s all cloud-native. How do we make it so it’s really easy to onboard these folks? They don’t have to change the way they do things, and we automate everything. That’s the beauty of DevOps. How do we automate everything, get it working? That’s the goal of this whole new team within the company itself.

Dave: [00:27:06] Really great point here, too, because those workflows have to change, and one of the things I found interesting working with the a lot of customers is there’s a big cultural impact, too, that has to change a lot to be successful with Kubernetes. I just found that really interesting how the culture and how you work needs to change, too, when you start adopting Kubernetes.

Bridget: How is that different from any of the other patterns that they were using with their config management or their VMs or whatever?

Dave: Yeah, I think the biggest thing is now you leave a lot of control up to Kubernetes, things that you had very tight control over. It’s no different than adopting a DevOps-type culture. A lot of those things are going to have to change in just how you operate and the controls you hand over to developers to empower them. Make them productive.

Lachlan: [00:28:08] I think the only thing I’d like to add is there’s no right answer about the best way to do this, but there’s a set of tools and techniques that we’ve all seen that have worked in different kinds of ecosystems and environments. This book goes into providing those sets of blueprints, which help shape the decisions you need to make when adopting this kind of technology or making cultural changes, which I think is incredibly valuable. Without this, you’re left scrounging around what is the best way to do, you know, push out an app and deploy and get developers onto things like Kubernetes or do network policy. This scopes down the touchpoints you would have to consider based on things that we’ve actually seen out there in the wild working for different customers or community members.

Bridget: Okay. We’re a little short on time, so what I would love to do is is get each of you to give us your best advice for best practices. Like, you wrote a book all about best practices, and, you know, I don’t know if you want to kind of touch on depending on context or how people evaluate and make the decision, but like, you know, I’ll just, you know, I don’t know, take you in reverse byline order. Let’s start with Let’s start with Lachie and say, what’s your best advice to people who are trying to have their Kubernetes best practices and read it too?

Lachlan: [00:29:35] Yeah. So for me, when I was out actually responsible for a platform that was built on Kubernetes, it was coming to the understanding that this is a journey that will take, you know, it will keep going. It’s not a destination. There will never be a point where I look at the whole system and say, it is perfect. I no longer need to touch it. So always adjusting and reevaluating the decisions you make. So best practices for me is around guiding you to make a decision because I also see a lot of people out there who are paralyzed by indecisiveness because there’s so many complex things that they can’t grok all at the same time. So taking these best practices and actually making a decision and moving forward, I think, is one thing that I would love people that read this book to, to be able to get out of it. The other thing is just you should be able to, no matter what stage of Kubernetes or your journey is with Kubernetes, reference this book throughout time and get something else out of it and continually reevaluate the things and the choices that you’ve made and make constant adjustments. Best Practices is something that you can always refer to, and it always gives you those nuggets of wisdom that you can take and course-correct or change out there in your running environments.

Eddie: [00:30:51] Excellent.

Bridget: And put a paper copy in your time machine.

Lachlan: In your DeLorean. You’ve got to put it in your DeLorean and make sure Biff doesn’t steal it from you.

Bridget: That’s right. Can you imagine? We thought the Sports Almanac was valuable, but if somebody had this book— All right, Dave.

Dave: Yeah, so the biggest thing I always stress is really, you have to walk before you run with Kubernetes, meaning that you really need to focus on those core concepts in Kubernetes and the core constructs that are available to you in Kubernetes before layering on a lot of complex technologies. I think we overcomplicate our environments. I don’t think Kubernetes has to be complicated. I think, a lot of times, we make it overengineered and complicated, because we have great tools like Helm, where I can give it one command and have a service mesh up and running. But there are implications to deploying these new complex technologies. Really focus on those core capabilities that are built into Kubernetes. Get really good at those, and then iterate and add in some more of these really useful technologies like service meshes and that.

Bridget: [00:32:06] Dave, I think I hear you killing my service mesh dreams, or at least saying you can’t start with that. Maybe start with a little bit more basic before you jump into really advanced.

Dave: Yeah, it may not be a great starting point for day one. Get good at those things like network security, resource management, policy. Those things are really important to be successful with Kubernetes.

Bridget: Nice. All right, Eddie, what do you got for us?

Eddie: Yeah, I think, obviously, echoing Lockie and Dave’s sentiments, definitely, but I think another one is When you approach best practices, especially around cloud-native and Kubernetes, if you’re coming from a long background and other ways of doing things, especially on-premises, old server VM processes that you did before, keep an open mind. The best practices you had there may no longer fit today, and really try to keep that open mind about how things work, especially when we are— I think Dave said it perfectly— giving a lot of control over to this new thing that’s That’s now kind of this data center operating system that’s managing our systems. Be open to just a new way of doing things because I think the biggest challenge that Dave and I have in the field is the bias of, no, that’s not how we did things. That’s not how we do things today. It’s like, okay, but this is how you’re going to have to do things. Then start layering. Let’s start from the simple and let’s work our way up to the complex, as Dave was mentioning.

Bridget: [00:33:42] All right, Brendan, tell us what you got.

Brendan: Yeah, I mean, I think a lot of it’s been covered. I think my biggest thing, just to echo, is to make sure that you understand why you’re making every decision and why you’re adopting every technology. Even Kubernetes itself, there should be a reason why you’re doing it, not just because it was in the CIO or CEO magazine that Everybody needs a Kubernetes strategy. You need to understand the system well enough before you even start to understand why you think it’s valuable to you. What is the pain you’re trying to improve on? I would say that for every single technology. I mean, I think that a lot of times people say, well, but I mean, in order to use Kubernetes, I need to have XYZ also. And you’re like, well, actually, why? People who are hosting simple websites behind an ingress and a service and a service mesh, and you’re like, I don’t think you needed all that stuff. And remembering that that all adds operational complexity, especially because I think Kubernetes makes some things really easy. It makes it super easy to deploy stuff, but it actually doesn’t help you understand the system. I think it used to be that systems were sort of proportionally hard to deploy, and so you learned along the way. I think one of the dangers that Kubernetes presents is that some things that are actually very hard to manage in production over time are very easy to get started with. People then assume that that ease of use is going to continue throughout the lifecycle of the tech, and it’s just not the case. Maybe another way of saying this is that Kubernetes is alive. It’s a living thing and not a static thing. That goes to Maki’s point about it being a journey. It could turn on you at any moment, right? Just because it’s working today doesn’t mean that you’re not going to have to investigate something and change something tomorrow. I think that’s a little different than something that’s been statically deployed to a VM under somebody’s desk and it’s been running for years. I think that’s an adjustment for people, too.

Bridget: [00:36:01] It almost seems like we have the Admiral Ackbar principle there, like, it’s a trap if you think that it’s going to be easy. But at least with this best practices book, it’ll be easier. Okay, so let’s bring it on home. Community events, stuff, where can our listeners catch up with you folks?

Dave: So, I will be at All Things Open talking about non-code contributions to open source, and also KubeCon North America, again talking about non-code contributions to Kubernetes.

Brendan: So, I’m going to be at the Kubernetes meetup in Heidelberg, Germany, coming up in a couple of weeks. I’ll also be at Microsoft Ignite. Orlando, if you want to go to Harry Potter World or Disneyland or Disney World, I guess. Bring the family. Yeah, no doubt, right? And I’ll be at KubeCon as well, KubeCon North America. And you can also obviously always hit me up on Twitter @brendandburns.

Eddie: [00:37:04] I’ll be— my next event, I guess, is my meetup here in Austin. So, I co-lead the Austin Kubernetes Meetup, so we have our October 24th meetup, so if you’re in the Austin area, please come on by. I will also be at KubeCon and Contributor Summit in San Diego in November. I may be at Ignite. I’m still deciding. Right now, I’m on paternity leave, so I don’t know what’s going on with my schedule just yet, but I’m back at work on the 14th, so I’ll know more then.

Lachlan: I will be at KubeCon North America in San Diego in November as well, so come and say hi. I’d love to chat about all things best practices or whatever your journeys are in Kubernetes. Come up and make yourself known. You can also— I also have a YouTube channel where I do OSS unboxings to help people understand all the tools out there and what they might do. So that’s a fun way to keep up to date. And you can always ping me on Twitter as well. So look forward to seeing everybody in KubeCon North America.

Bridget: [00:38:04] Awesome. And I will be Let’s see, I’ll be at Twin Cities Startup Week in a couple of weeks, and then DevOps Days Philly, and then DevOps Days Ghent, and Velocity Berlin, and KubeCon, of course. In terms of something fun to check out that totally isn’t tech, theoretically today, they’re out for delivery. Joe and I are getting some VanMoof e-bikes. I am very, very excited about those. Apparently, the correct number of bicycles to have in your life is n+ 1. This may not be the end to everything, but we now have 6 bicycles. Well, when those come, we’ll have 6 bicycles, 2 for me, 4 for him. Then we also have links in the show notes to all the things people talked about, and the OpenCFPs for DevOps Days are devopsdays.org. And Velocity Berlin, there’s, as well as many DevOps Days, the discount code ADO2019 will give you discounts, 20% off in many cases. If you head over to arresteddevops.com/kubernetes-best-practices, you’ll have this episode’s show notes. Visit arresteddevops.com/itunes, leave us a review in the iTunes Store if you want to help other people find the podcast. That is apparently a thing that exists in the world. I have no idea how that works. We’re also apparently on Spotify and iHeartRadio now, if you’re into those systems. Thanks so much to Brendan and Eddie and Dave and Lockie for joining today.

Brendan: [00:39:39] Thanks, Bridget. Great to be here.

Bridget: I’m Bridget at Bridget Kromhout. This is Arrested DevOps, and remember, there’s always DevOps at the banana stand.

BROUGHT TO YOU BY

Bridget chats with the authors of Kubernetes Best Practices: Brendan Burns, Eddie Villalba, Dave Strebel, and Lachlan Evenson

Bridget talks with all four authors of the forthcoming O’Reilly book Kubernetes Best Practices: Brendan Burns, Eddie Villalba, Dave Strebel and Lachlan Evenson. Brendan writes because “I like to teach,” a legacy of once being a professor. Eddie has been at Microsoft for 10 years and wants to spread what the big organizations learn, good and bad, to startups that can’t get that help. Dave helps customers succeed with Kubernetes daily and never aspired to write a book, but likes breaking down complex technology. Lachlan wanted to give back to the community, remembering Brendan standing in a hallway in late 2014 or early 2015 answering all of Lachlan’s questions, and to write the book Lachlan wished existed in 2015. The cold open is Brendan’s line: “It’s a powerful tool, but it’s also kind of a footgun.”

Short Essays, Not a Narrative

Brendan says the project has moved from something people heard about to something everyone wants to implement, but people struggle with specific tasks, and hands-on help doesn’t scale. They’ve seen “lots of people sort of shoot themselves in their foot.” Unlike general introductions to Kubernetes, the book is focused on specific topics, to dip into when working on machine learning or setting up a cluster for a bunch of developers, so it’s “a series of short essays rather than a whole put-together book.” The 258-page PDF Bridget has in front of them has no narrative flow, which is on purpose.

Lachlan says now is the right time because adoption has grown, the ecosystem has become more complex, and Kubernetes has a sprawling variety of APIs, so the book shows where to start on topics like policy, rolling upgrades, governance and security. Eddie says organizations are already down the path and don’t want another step-by-step walkthrough, and the authors tried not to make it a snapshot of one version. Dave says users need to focus on the core concepts and often skip them to over-engineer. Lachlan likes the mix of philosophy, meaning why you’d want policy, and the tactical how.

Chapter Favorites

Lachlan’s favorite was Chapter 11, on policy. Lachlan notes “everybody loves hearing Chapter 11 for anything,” and says enterprises moving workloads to Kubernetes ask how to make sure workloads conform to policy, whether regulated or just wanting to understand configuration. Bridget notes the chapter covers the open source project Gatekeeper, and Lachlan says it’s a Kubernetes-native implementation of OPA, the Open Policy Agent.

Dave’s was resource management, which “doesn’t sound really exciting at all” but is something users struggle with and affects scaling. The book covers best practices around requests and limits and how workloads behave when capacity runs out. Lachlan says most of the outages Lachlan was paid to handle in the early days came from resource management, as clusters got to 80, 90, 100%, and would have liked to have had the chapter in 2015, to avoid a cluster going into cascading failure at 3:00 AM.

Eddie’s was Chapter 9, covering networking, network security and service meshes, which was the most challenging to fit into a concise format. Eddie calls networking the foundation, where little things trip people up, like the move from kube-dns or SkyDNS to CoreDNS, and describes a customer where divisions put Kubernetes in without telling anyone, and then security asked why their controls were gone. Eddie says people want a service mesh as an “easy button” for observability, security and policy, and find it’s “Thousands of little buttons that you have to press in the right combination.” The chapter describes what all service meshes should do, what to prioritize, and the SMI spec, a common API for those things.

Brendan’s favorite, after the first chapter on laying out a service, is the one on developer workflows. Brendan worries that operators love the technology, or it’s great for continuous delivery, “but we’ve made the developers’ lives miserable.” The chapter covers partitioning a cluster with namespaces, onboarding a new developer, RBAC so people don’t step on each other, cluster-level logging and monitoring that’s just there, and testing and debugging, since “if it’s not easy, people will do less of it, and then you ship buggier software.”

Will Organizations Do This?

Bridget asks whether people in the field actually follow these recommendations. Brendan says people want to, and “the recipes just aren’t there, necessarily.” Eddie describes a customer creating a new division to build patterns for developers and ops so onboarding is easy and everything is automated. Dave says there’s a big cultural impact, since you leave control to Kubernetes that you used to hold tightly, much like adopting a DevOps culture. Lachlan says there’s no single right answer, but a set of tools and techniques that have worked, which the book offers as blueprints, so you aren’t “left scrounging.”

Best Advice

Lachlan’s advice: this is a journey and not a destination, there will never be a point where the system is perfect, and people paralyzed by indecision should use best practices to make a decision and keep adjusting. Dave’s is to walk before you run, and get good at the core capabilities, network security, resource management and policy before layering on tools like a service mesh installed with one Helm command. Eddie’s is to keep an open mind, since practices from on-premises servers and VMs may no longer fit, and the bias of “that’s not how we did things” is the biggest challenge in the field.

Brendan’s is to understand why you’re making every decision, including Kubernetes itself, and not adopt it because “Everybody needs a Kubernetes strategy” appeared in a magazine. Brendan says Kubernetes makes deploying easy without helping you understand the system, and that things hard to manage over time are easy to start, so people assume the ease will continue. Brendan’s summary: Kubernetes “is alive,” a living thing and not a static one, that “could turn on you at any moment.” Bridget calls it the Admiral Ackbar principle: “it’s a trap if you think that it’s going to be easy.” The episode closes with a joke about putting a paper copy in a time machine, or a DeLorean, so Biff can’t steal it.

Bridget chats with the authors of Kubernetes Best Practices: Brendan Burns, Eddie Villalba, Dave Strebel, and Lachlan Evenson. (Kindle version available now!)

Community & Events & Stuff

Dave:

Brendan:

Eddie:

Lachie:

Bridget:

If you have an upcoming conference you would like to see promoted on ADO, you can fill out the handy form at arresteddevops.com/conf

Upcoming conferences

Open CFPs

Discount codes

This episode's guests