Now Playing

2025 EES The Tech Hustle: Behind the Scenes of AI at Scale

EES 2025 The Tech Hustle Applied AI Engineering
Details
Length
57:10
Views
9
Featured Company
The Tech Hustle
Published
Oct 8, 2025
Find Your Path Forward

Ready to build your tech career?

AI-native engineering courses taught by senior engineers from the companies you want to work at. CS students and CodePath alumni enroll at no cost.

About this video

The Tech Hustle takes students behind the scenes of AI at scale: the systems, the teams and the trade-offs that show up once a model has to serve real users. Moderated by CodePath’s Dr. Anthony Davis at the 2025 Emerging Engineers Summit.

Transcript

Hello. Hello everybody. I am Dr. Anthony Davis.

I’m the senior manager for development and engagement here with Copath and I am excited to be helping to moderate this session. So so we’re excited to have you. So before I start the next session we I just want to introduce a few things.

So we just want to remind you all to make sure you’re utilizing the chat for your comments all of making sure you put everything within the chat. Okay. Please take a moment at the end of the session to give us to give us your feedback as well.

And and from there I’m just excited to introduce this next session. So the next session is going to be called beyond the scenes of AI at scale with Bobby D. Okay.

And let me give you a little background about Bobby D. Bobby D is a seasoned in engineer and professional speaker with 20 plus years of experience in technology and engineering specializing in the practice of site reliability engineering with the emphasis on hyperscale distribution infrastructure. Bobby DD is known for his outgoingness and historic knowledge of Twitter tech and DNI journey where he worked for nine years becoming one of the only few black engineers to achieve the title as of staff engineer.

As Bobby D journey continues, he looks forward to the opportunity to continue to continue contributing to his community hash the techhustle by mentoring, coaching, and teaching the next generation of tech and the value of our diverse perspectives. So, let’s give Bobby a a round of applause as we bring him to the stage. What’s up?

What’s up, everybody? Anthony, thank you so much for that introduction. I really really hope everyone has had a great few days or a few hours getting EES started and I can’t imagine how the next few days will be.

Today’s conversation and I believe it’s title, let me see what I wrote down here. Sometimes I’ll be forgetting it says behind the scenes of AI at scale and we’re basically going to be really talking about tech today. I’m one thing that I always enjoy doing, especially with mentees that are getting in introduced to the industry or may not have attended a real hardcore tech talk.

I got one for you today. Now, don’t beat yourself up if there’s any technologies or any terminologies that I may use that might be a little bit over your head. It’s all right, but it definitely gives you an opportunity to let you know of what skills you can close in terms of gap.

But also why I’m in these streets, you know, as a a mentor, a coach, and definitely somebody that’s been around the block a few times. And we’re going to jump into the slide deck and definitely get you all a chance to to see what I’ve been talking about. Let me see if I can click the button here.

Let’s see. All right, perfect. Yes, it looks like my slide is on screen now.

And yeah, we’ll jump right into it. We are going to be talking about behind the scenes of AI at scale from from distributed systems or distributed inference to system reliability. And it’s presented by me Bobby Dorless.

Now I know some of you may have met me today earlier. Big shout out to y’all who attended the keynote speech. And definitely the opener and y’all probably going to see me on camera throughout the next three days.

But I want to give you all a little bit more formal introduction of myself, talk a little bit about my journey before we jump into the conversation here. And I definitely want you all to use the chat room to throw questions out there. You know, the team is here to support us to to, you know, get me questions in front of me if there are.

But definitely closer to the end of our conversation, we’ll definitely have some some chances for more questions. But a little bit about me. As mentioned, my name is Bobby Dorless.

I’m a technologist. I’m a conscious engineer. I’ve been at the keyboard for more than 20 plus years.

I know sometimes it sounds a little intimidating and I like to make fun of it because I do want to let you know like at the end of the day this is not new to me, right? I’ve been around the block a few times and definitely have seen technology iterate from back in 2002 when I first started to where we are now in 2025 and I tell you my mind is blown in terms of where we’re at in terms of the iteration and progression. I’m the founder and CEO of a tech community called the Tech Hustle.

Where our goal is primarily focusing on mentoring and coaching the next generation of engineers by using different mediums like featuring on platforms and organizations and companies like CodePath to running a podcast where we have a podcast where we speak weekly about technology engineering. And if you haven’t heard yet, we run a really cool newsletter that postes publishes monthly talking about tech really from a different perspective. Not going to give you no sales pitch here, but definitely want to encourage you all to link up with me on LinkedIn.

Find our website www.thetechhustle.com and get registered for this free information that we’re giving out there because one day I might start charging for it. Now, in terms of my background in technology and engineering, I mentioned that I’ve been in the field of engineering for 20 plus years. I started off at the help desk.

So, my journey into tech didn’t come from, you know, attending a a school in college-wise or getting an internship and then eventually getting a really cool job. But it actually started off working at, you know, the the entry level position in terms of technology being the help desk. I did attend a little bit of schooling where I only had two years of technical school and use certifications as a way to level up.

But I really enjoyed my journey because I’ve really seen from, you know, in the mud all the way to working at, you know, a famous well-known social media company called Twitter. At Twitter, I became one of the lead engineers for the compute platform there. And this is why this talk right now is so dear to my heart because designing and building computing infrastructure at Twitter scale is and sometimes I always tell my mentees it’s like one in a billion chance that you were able to see it being done and let alone these hands actually built it.

So I know what scale really looks like because the team that I managed for almost you know five years I was there for almost 10 years 2013 to about 2022 right before Elon Musk got here. So, I know a lot of people ask questions about that, so I don’t really know about what happened after Elon Musk. But during my tenure there, I really foundationally worked on teams that build data centers, build infrastructure, and actually build the platform that 90% of all of Twitter ran on.

Now, I may have mentioned some technologies out the gate already where I said data centers. Data centers are basically these big old warehouse buildings and we stack computers from the floor to the ceiling. And basically somebody needed to manage and design systems and infrastructure that was going to support that and that’s what my team specialized in.

My skill set is in software engineering but I’m also a systems engineer is actually where my most proficiency at. I’m a Red Hat certified engineer. I know from you know inside of a computer how all that stuff worked to connecting it to the internet to eventually installing apps and making the world depend on it right in terms of reliability and more specifically my title at Twitter when I left was a senior staff site reliability engineer if you’re familiar with site reliability engineer it definitely revolves around designing and maintaining the reliability of a site and the site that I was responsible for was Twitter which sometimes after I look back really blows my mind in terms of the the scale that we’re at and actually how much the world depended on it and a lot of people didn’t know that it was me back there, you know, taking care of different things.

So much so that this is the reason why I’m passionately into talking about technology at a really high level. Because number one, you don’t see that many people that come from a community that looks like me and you that have worked at that level and can come back and talk about it, right? And that’s what I’m really passionate about.

My mission as mentioned here is that I’m really trying to transform the tech landscape by empowering the next generation of engineers because there’s this wave and I mentioned this in another talk that I did on stage a few weeks ago. There’s this wave that’s only two letters and I know everybody knows it’s it’s called AI and the cool thing about waves are if you turn your back on a wave, right? You you’ll get hit and knocked right over, right?

I’m from Florida, so we go to the beach and that that was like number one rule. Don’t turn your back on a wave. What you should do is position yourself, line yourself up for a wave and get ready to jump when the wave comes, right?

And that’s what I feel like AI in terms of what I’m going to be talking about this evening. Is how we’re designing and building infrastructure to support this wave that’s approaching. And it’s and it’s truly is cresting at this point.

So, I do want to encourage you all to dial in. Like I said, there will be some topics that I’ll talk about that might be a little bit over your head. That’s cool.

Write them down so you can go check out some, you know, terminologies or, some specific skills that we’re talking about. But the main thing here that we’re talking about today is what is the challenge with AI at scale. Now, I know most of y’all have heard AI.

I know most of y’all have used AI at some capacity. If you haven’t used AI yet, I’m going be straight up. This might not be the industry for you because if you’re not already using AI right now, you’re already getting late left behind.

And I and I mean this wholeheartedly, not to throw any shade, not to, you know, downplay your progression and where you’re at in your journey. I just want to encourage you to use it every single day if possible. But there is a need to build computing infrastructure to support the scale that all of us technologists are talking about where we’re going to be going towards and the challenges that we run into in terms of scale is in those buildings that we talked about as data centers there’s limited of space right how many computers can we get into that one building another challenge is is how much cooling like if we stacking computers from the floor to the ceiling how are we going to cool all those computers I You’ve walked around before with your phone, which I don’t know if you knew is a computer in your pocket.

Yeah. Yeah. Yeah.

But if your phone is overheating or you left your camera on, right, and all of a sudden on the side of your pocket or you may have it on the table, you feel your phone and it’s warming up. It’s because it’s it’s heating up. Now, inside of data centers, these racks of computers that’s depictked right here in this illustration creates a massive amount of heat.

But you need to build a system and infrastructure that supports the heating and cooling of computers especially when we’re talking about GPUs. Now the other side of that is not even just the building and we’re talking about cooling but the demand for power. I don’t know if you all knew but the amount of computing power that is needed for us to take the next leap in AI we don’t have it.

I read a paper the other day that said if every AI lab turn on their AI, the the United States power grid will just crumble. Especially with all the projections of how many data centers and, you know, new projects that are being kicked off. And I don’t know if you’ve seen some of these ticket prices.

They’re talking 500 million. That’s half a trillion dollars. I I apologize.

500 Billion. That’s half a trillion dollars of investment that all of these companies are making in building out this infrastructure. That’s why this talk like this in terms of talking about AI is scale at scale especially now you’re getting into your career is going to be pivotable in terms of your iteration of where you’re going to be going in terms of direction.

Now computing as a demand right it is a multi-billion dollar industry. Low-key, I feel like AWS, GCP, and Azure or whatever m Microsoft calls it Asher or Azure is is the new hustle, right? A lot of companies have already move to the cloud, but the cloud is actually real computing resources.

I I’ll tell you a quick story. I remember I was talking to a mentee of mine about what the cloud is, right? And it was a younger student that was really interested in tech but couldn’t really find where their lane was and we started to talk about cloud engineering and we started to really dive into what is the actual cloud right and this may be something that you have pondered right hey I I have computing resources I have my emails I have this database that I can access and it’s out in the cloud somewhere but what is that and there’s this thing that I always want to remind you we are communicating from one computer to the next computer.

When your phone connects to the internet and it connects to the cloud or connects to a service that’s running in the cloud, there’s a physical computer that is running that process. Even though it’s virtualized, even though it’s containerized, there still is a physical computer that you have to talk to on the other end. Now, who manages, who develops, who designs, who supports those computing resources?

A lot of companies, as I mentioned, decide to move to the cloud because, you know, managing computers is a tough job, right? And even as a an attestant to this is how many students are actually studying, you know, systems engineering in the chat. Throw that in there.

Like if you are studying systems engineer, put a plus one. But mo most students or most CS students especially are studying more on the computer science software engineering side of the house but there’s this other practice around computing resources that out there and that’s why talks like this you know uplifts us on the systems engineering side of the house but also gives us an insights into the value around our skill set I’m going to say that again the value around systems engineering skill set if they’re talking a lot about software engineering and AI having a battle there’s still going to be a need for systems engineers to build infrastructure at scale. But the other part about compute, right, and this is something that I always learned while working at Twitter was there’s always two sides to the coin.

One is that you build infrastructure that’s going to be managing compute. And let me just re roll back in terms of what I mean by compute. Compute is when I’m talking about CPU, memory, and computing resources.

Right? If you write a program that you’re going to schedule in some type of cloud infrastructure, you need computing resources like CPU, memory, disk and network for you to run that process. But you also need a place to store your data, right?

And that’s why data and compute are always considered as two different entities, right? And that’s why when I break it down into the challenges with scale, they’re actually two separate challenges. And even on the compute side, I forgot to mention a technology that I hope you all get familiarized with, which is called GPUs.

It’s another type of processor that we usually utilize in graphical processing, especially with the, you know, number one hardware provider being Nvidia. We used to put these graphic cards in home computers so that you could play your video games better. And now these same graphic cards are being used to train but also you know distribute inference for AI.

It blows my mind that that computing technology has gained so much popularity but that resolves in the compute platform. Now on the data side of the house the most important thing about data is number one is accessibility in terms of how can I access the data that I need how can I access it reliably and how can I access it in a way that doesn’t hinder me from processing whatever I’m doing right either I’m storing it inside of a database right so that’s more like less like a postgress or my SQL you also have no SQL type databases like MongoDB B do I store it in let’s say even something more simpler than that in a in a raw file called SQLite. SQLite’s a great example of storing data in a specialized way that makes it easily accessible to anyone.

And then you can even go deeper into that where you’re talking about object stores, right? So it’s like, hey, as and we’ll use Twitter as an example. As all these tweets are coming in, we need to store all of these tweets inside of an object store or in some type in some cases a key value store where we say hey this is the key for the the tweet and then the results in terms of the data is actually what’s inside of the the tweet itself.

Object store and blobs are going to be like images like whenever you upload an image to let’s say Instagram or Facebook or even Twitter they’re being stored in an object store system and a blob store system u that allows us to be able to serve that back but if you notice something you can see that these two entities are treated as two separate things right and the thing about data and if I would have known this even more as I was coming up in the industry is how valuable data is I actually tell my mentees that data is by far one of the most valuable assets that companies create. And sidebar, the reason why Elon paid $44 billion for Twitter was the amount of data we had about everything. And I mean, we had tweets from the first tweet.

And yeah, we deleted tweets, but we’ve always maintained consistent copies of tweets replicated all around the world. Now, that data, how valuable is that data when you’re training AI? Do you see the connections?

Like being able to maintain and support data for long-term storage, either being hot or cold storage, but also utilizing data in a way that we didn’t even know that we were going to be utilizing it to train AI models is massive. And if I recall when I left, Twitter had hundreds of pabytes of data, that I know for sure were being used to build out different type of models like whatever XAI is built and also other, you know, AI labs that actually had access to backend data at Twitter. And overall, the thing that I want to, you know, emphasize here, especially about challenges of scale, is that we’re not dealing with gigabytes.

We’re not dealing with terabytes. We’re definitely dealing with pabytes of data that you have to design systems so that you can scale as you’re growing but also utilizing them. And then reliability requirements, right?

I think that right now AI in its current form is really taking advantage of what we’ve built in our industry in terms of reliable services, right? The simplified interface to access OpenAI is an API call, right? Application programmable interface that relies on a restful API that’s in a JSON type data structure that if you have a access key or token well I would say more of an access key an API key you can access the API and then you receive data back from an API call in JSON for I mean AI was basically primed to be in a more reliable state because of a lot of the technologies that we built out throughout over the years.

And definitely something that I lean on as a site reliable engineer on a daily basis is definitely the reliability of systems. Now, when you’re talking about what hypers scale architecture looks like, we talked about compute kind of dives a little deeper into, you know, past experience at Twitter and really, you know, visualizing what the cloud is. Also a new computing resource called GPUs.

And then you have storage systems that actually store actual objects that are utilized by AI for training but also for just storing data in different type of formats. But when you think about what scale looks like and this this is illustration right here that really gives a quick depiction of how clients that are all the way on the lefth hand side the little small client accesses open AI which is all the way in the back right is there is a lot of opportunities for us to design system and architect them in a way that sometimes make it look even more simpler than what it really is or sometime it’s a little too complicated. And then that’s why users have a a kind of you know unpredictable experience.

The one thing that I always like to teach especially when I was mentoring at Twitter is keep it stupid and simple that anybody can just come in and pick it up and keep it running. And that also applies to when you’re designing systems. Now, I am using this word system design, designing systems overly obsessive right now because whenever I hear a mentee talk about they’re about to go to an interview and they have this interview called system designs, this is what we’re talking about is what does the infrastructure look like when you’re designing the flow from a client to an AI or an API call on the backend system, right?

But all that’s in between, especially at the compute layer, also includes the different type of hardware that you potentially would have access to. We’d already talked about GPUs. There’s something called TPUs.

I believe Google’s created that and have a more proprietary buildout of hardware. And actually, I just read something today that OpenAI is partnering with, I believe, Broadcom or some other, you know, company that makes chips and open Open AI is going to build their own chips. Right?

This is where the industry is moving. Like it’s like, hey, we don’t need anybody else but ourselves. And they’re making real large big investments into building out these computing resources.

But when you’re talk thinking about architectural wise, let’s take a look at this diagram here. You’ll see that the client talks to something called the edge and the edge kind of is the kind of middle person between cloud and then the cloud being able to talk to the backend system being open AI. Now, when we’re talking about the edge, this right here is something that I think a lot of users don’t realize how we trick you all on how something called latency is actually achieved.

If you’re familiar with latency, latency is the time it takes for a packet to go from my computer to the other computer and back, right? And we want everything in under 200 milliseconds. If you’re not familiar with 200 milliseconds, when you get a chance, go on Google, go on ChatGPT and say, "What’s the importance of 200 milliseconds to a software engineer?" And you’re going to get a plethora of information around what 200 milliseconds means.

But the 200 milliseconds means that we have to make all this happen before the clock runs out, which is 200 milliseconds. And when you move things closer to the edge, what does that do? It basically reduces latency.

Now, sometimes you’ll have to go from the edge to the cloud or into a physical data center, but your first line of defense to reduce latency is always moving stuff closer to the edge and then go home if you have to, right? And then going home obviously has to be designed in a way that allows for slow I’m sorry, for fast response and lower latency also. And then that’s where a lot of these distributed systems come into play such as distributed file systems, object stores, and more or less the way that you design storage systems to react.

And why do I jump from compute to storage? If any processing happens in compute, it happens the fastest. Any processing that happens in compute and also attached to physical memory like memory inside of your virtual physical box or your your container, everything is fast.

Soon as you have to exit out of the CPU cycle memory that’s local and you have to go and fetch things across the wire for data. This is usually done through a let’s say a SQL call or let’s say Postgress MySQL key value store whatever you’re increasing latency that’s why storage in terms of comparison is always honored in terms of how you have to design storage systems so that you reduce latency. One of the coolest storage systems I ever worked on, and this is another sidebar, is something called Reddus.

Reddus is a database engine that’s written in memory. So now imagine you’re accessing a a storage system where you have to leave your CPU in memory and then go to another storage system and it its data is stored in memory. Obviously, that’s going to make it response even faster, right?

We use Reddit for a lot of caching. Usually have it on the edge of the internet, but you find backend systems taking advantage of some of those really cool ways of storing data in memory. But also, this is where algorithms really come into play in terms of how you’re storing items.

Let’s say from an array to a set, right? You definitely have two different bigo notations. If you’re writing to an array, O of one and then of set a set is more or less being able to respond even faster, right?

But in general these are the differences that you have to consider from compute to storage and then an orchestration layer. Now what I what I mean by orchestration layer one of the things that a lot of people don’t realize is the amount of effort that it takes to manage and build out this infrastructure. Right?

I’m talking a lot of terms here compute disk storage that but how does this all tie together? Right? This is where the orchestration layer comes into place where you’re really designing systems to be more how can I say very generic and when you’re scheduling computing resources and what I mean by this I talked about a tech earlier I said something like virtualization I also mentioned something like containerization but there’s this platform that sits right above all of that that basically has for better or worse become the more standard way of creating a more orchestration layer and it’s called Kubernetes.

Not sure how many of y’all know about Kubernetes and definitely want to you know encourage it as something the skill set to to to develop. As a software engineer you’re basically interfacing with Kubernetes to schedule your workload, but as a systems engineer you’re building the infrastructure that’s going to run Kubernetes. Some really cool stuff, right?

But Kubernetes as an orchestration platform is only one example. And it’s definitely something that I’ve grown and really appreciated over the years because while I worked at Twitter, this is another little sidebar. We built our own orchestration system.

So, I don’t know if you all knew this, but even with Kubernetes, and the system that we built at Twitter called MSOS and Aurora and what Facebook runs, they call they run something called Tuppawware is all foundationally based off of a platform created by Google called Borg. It’s a great research paper to read up on. Definitely want to encourage you all to take a look at what it looks like to design a system to be hypers scale and definitely fall in line with what we’re talking about.

But Borg itself is basically the the the beginning of where we change the concept of computers being individual objects themselves and looking at that big data center as one big computer. I know it it it doesn’t make sense but just imagine when you’re managing computers individually right that creates number one toil and obviously you need to increase headcount because you’re create you’re managing each computer individually but when you put it under an umbrella like an orchestration system like Borg what Google created and then one of the you know the the the new products that came out of Borg which is Kubernetes it allows you to look at your infrastructure in a different way where you’re looking at the whole building as one computer and then you’re looking at applications where you’ll just schedule those applications anywhere in the data center right it’s some really cool tech and I definitely want to encourage you all to double tap on what orchestration looks like now even with orchestration systems storage systems and obviously we talked about compute we’re talking about AI here right and I’m not going to actually I did mention their name earlier but I’m I’m not going to give you any financial advice but if you are keeping a track on the value around the company that is the primary developer of GPUs. You can quickly understand why I’m going to stand on this is that this technology, this hardware technology called GPUs will be the future of what computing looks like.

And we need to understand how GPUs differ from CPUs, from the way that they process memory, to parallel processing, and more or less how there might be other hardware accelerants out there that could accommodate some of the similar use cases of GPUs, but more specifically, GPUs as an optimization technique is something that I really want to encourage you all to dive a little deeper into. And remember, I know most of you might be at the software engineering side of the house. There might be a small few of you at the systems engineer, but there’s definitely a gap in the bottom called hardware engineering that I want you all just to spend some time maybe read one or two research paper, follow along and be familiarized with this technology because this tech influences the stacks above it.

Right? So if we’re designing a system or designing a a a hyperscaled AI system and a software engineer doesn’t understand what computing resources we’re allocating, they may be not they may be designing the wrong type of system to run on that platform and it’s happened before, right? And that’s where the gap and in my opinion always need to be closed and opportunities like this for software engineers to learn the low-level infrastructure things is definitely one of those chances to be very inspirational, right?

So the other thing that I want to make sure I mention in terms of scale itself is what distributed infra inference systems look like. So when you’re talking about interacting with an API this is definitely something that I always like to quiz you know potential candidates is what happens when you type in www.twitter.com twitter.com and you hit the button enter and you traverse the internet to get access to whatever let’s say backend data that you’re accessing through an API. That same model is what plays with us in terms of our interface today and it’s a model that leverages some already existing technology like load balancing and then you see me talk about latency optimization and more or less what we’re talking about is the model service architect.

The model service architecture looks very similar to restful applications or back-end systems. I’m not sure how many of you all are on the backend side of the house. Me my most proficient programming language is Python and Go.

I usually write, you know, if I’m writing a backend system, I’m using Flask or Fast API. And I know a lot of you are out there that, you know, write and utilize Java and or Go or even just Node.js as a backend system. But all of those skill sets apply to what’s going to be coming up in AI.

And that’s why I bring it up is that a lot of the conversations that you’ll hear a lot of people talking about AI is the model building. This is where they’re building those massive gigawatt data centers where they’re just going to run, you know, the build process that takes, you know, six months to get through. What I’m talking about is the results of that.

So after they run the build process, there’s this binaries, these sets of binaries that basically is the binary of the AI, but who’s going to serve this binary? This restful type of API system that you’ll access and what we access right now. And that’s why I more or less want to just emphasize that this technology, this modern way of interacting with data doesn’t change as much with AI.

It’s just a different back-end system that or what I mean by backend system I mean like when you build an API you build an API and then you have a database right but everybody interacts with the API the database can change but the API doesn’t change that same concept is what’s being developed with open a with AI and open AI definitely utilizing it load balancing strategies now we can talk for hours about load balancing and definitely want to encourage you all to familiarize yourself with this technology around roundroin let’s say there’s different technologies like engine X haroxy but these type of architectural resources and tools and more specifically models are not new so that’s why I’m always encouraged especially if you all are developing the skill sets is that there’s a lot of resources out there to dive even more deeper into it all right now in terms of what reliability engineering for AI looks like right I started off earlier talking about while working at Twitter I was a site reliability engineer actually right what his name is site reliability and yeah the sites reliability dependent on my group of engineers and the same thing comes into play for AI right is that you’re designing systems that are going to be fault tolerant and gracefully handle failures right when we talked about load balancing if you’re not familiar with load balancing is basically like a traffic cop when a packet comes in through the front door the load balancer receives it after it goes to the firewall the load balance receives it and the load balancer says all right you need to go this way you need to go that way you need to go this Right? And it’s more or less optimized for failure because if one of those instances that it’s going to throw the packet to doesn’t exist no more, load balancer says, "All right, I won’t send no more packets there. I’ll just send it somewhere else." That’s what fault tolerant designs and and those type of strategies really come into play.

Monitoring and observability. I have a saying with my mentees that if you have a service that you do not monitor, then you need to assume that it’s broken. If you’re not collecting metrics on a service, then it’s broken, right?

Is monitoring and observability of the performance of not just the overall infrastructure but down into the instance of applications or instance of AI. It’s going to be you know a needed feature or skill set that you should develop. And then incident response.

This is really what site reliability engineers really focus on is how do we not do the same or do a you know how do we not have the same problem more than once right I used to have a saying if if something failed the first time it’s okay but if it failed the second time then it’s all our problem type thing right and you definitely want to make sure that you have processes in place especially when you have incidents with AI and I’ve been following along with open AAI each outages that they have I’m subscribed to their status page and they are definitely showcasing what it means to run systems at scale but also at a production level because you have these normal practices in place like incident management. Now, what does the real world look like and use case for AI? And we’re going to quickly skim through this because we’ve already kind of talked about Google, Microsoft, and a few other cloud providers out there, but just know that this technology this infrastructure, this scale is is going to continue to grow.

And there’s still going to be a need for the skill sets around for companies like the ones I listed here to really get through the process itself, right? So definitely want to encourage you all to follow along these companies in terms of what their real world use cases are. Especially Google there they I I mean if I go back two two years ago maybe even less than a year ago I thought they were lost in this fight here or this you know competition around AI but as of late they’ve been killing it.

And I knew they had the computing resources and I also knew that they had the engineering capabilities to really take us to the next leap. Obviously they were the ones that wrote the, you know, the transformer paper on how to teach and train AI. So, I knew that they were going to catch momentum.

But it’s always something that I want to encourage you all to follow along, read up on these, research papers, use case, and understanding the technology below. Especially when we get to like virtualization and containerization because that’s that’s how we actually manage and run at scale. And the future trends things that I really want you all to keep an eye out for as you’re going through your journey.

Especially as you are you know getting familiar with how this industry shifts and change is that you’ll notice that hardware associated with AI is only going to increase right there’s a demand for us to achieve a certain level of AI and you can’t get there unless there’s innovation in the hardware spectrum or the hard hardware side of the house so if you are electrical engineer mechanical engineer andor like me that likes to break things right is definitely want to give you some words encouragement that in terms of pathways to supporting scale and growing an AI scale is definitely building up efficiency around that. The other one that I mentioned here is edge computing or edge cloud hyper or hybrid computing. There’s this quote that I actually been you know mentioning to a lot of friends of mines about where the edge of the internet is.

And let’s just get some clarity first. What is the edge of the internet? First of all, think of the internet as if it was a highway.

And literally it is a highway. There’s wires that run under the sea across the country, across the world to connect us all on this highway, right? Some of it is being beamed out to space.

Obviously, Elon Musk got some, you know, satellites out there, but majority of our internet traffic happens on the planet and physical fiber optic cables, but those fiber optic cables have centralized points that they connect to. Just like a highway, when it gets to a city, that intersection or that junction is considered the edge of the internet. And what happens is a lot of companies put computing resources right when the edge of the internet intersects right there’s certain points and places in the United States that you know you’ll find more data centers being developed because that’s where the edge is and the closer you are to the edge the faster latency you have right so much so that I also have this feeling that the cloud providers are going to be the new edge of the internet and what I mean by that is everybody’s scheduling everything out in the cloud why do I have to leave that building to go somewhere somewhere else to be on the edge per se.

I should just stay inside those building or stay inside that region and I’m on the edge of the cloud edge environment and that’s why just talking about edge I think over time it would evolve and that that infrastructure in terms of scale is going to have a high demand for AI and think about just sustainability we talked about challenges of data centers cooling power and all of those difficult conversations I’ve had in the past when we build a building out there we filled up with computers and we’re about to go to production and then the facility teams come up and tell us like hey we can’t keep this building cool enough so we can’t right it’s like taking those things into consideration when you are building scale is always something to to to put on the table and definitely something that you want to plan in the future but hey we were able to get get around some of those challenges and then we’re coming really close to the end of our talk here and definitely want to cue you all up for it if there are some questions I believe Anthony will come back up or throw them in the question screen here for for you all to jump in. But some of the key takeaways that I want you all to know about scaling AI and understanding this you know wave of impact but also an opportunity for you all to develop is architecture matters right hypers scale AI systems requires like thoughtful multi-layer planning and architectural skills it’s not something you get just out of school it’s something you build over time that’s why when I hear candidates talk about the system design interview it’s usually something we reserve for a more, you know, a level two engineer or someone moving into a senior staff level, right? But you also have an a deep respect for architectural design because it does matter, especially at the beginning.

And then we’re always trying to optimize or being optimized with the computing resources and and definitely taking consideration the future development of technology, right? I know the GPUs are the hottest things on the street right now. I definitely want to encourage you all to get familiarized with the use case.

But definitely don’t want you all to lose out on where we’re going in terms of other tech. And then reliability designs. Just know that the the design of your infrastructure is really and should be foundationally principled on a few things.

S sur practices is by far and in my opinion obviously because I was an S sur and I really support the movement is definitely one of those primary kind of say means of developing a reliable system and the reason why I say this and I’ll take a quick sidebar on this. The reason why I’m so like shouting from the mountain top around site reliability engineering. I don’t know if y’all were around maybe you know 10 15 years ago when everything used to break from Facebook from Instagram Twitter used to break all the time Yahoo used to break AOL all of these systems used to break all the time until a team really put together different practices that got us all aligned in where we should be focusing our time on reliability and site reliability came out of Google.

So, big shout outs to Google again. But this practice itself is something that I do want to encourage you all to familiarize yourself because when you get to the industry, when you get into your full-time job, the the most scariest thing for me is when I break something or when something’s broken and I don’t know how to fix it. And the thing that you have with SRRES as a practice is different ways that you design systems.

So, you don’t get that feeling, right? That feeling like, oh, am I if I deploy this, will it actually work? Andor if everything breaks, what do I do?

Right? But I always do want to encourage the that that this idea around S sur principles is something considered in design strategies that you all will be exposed to. So I do know that we have I guess we have about eight nine more minutes left on the clock.

So if there are any questions I’m going to click the question bank here and see what we have here. I’m going to see if I jump into some of these questions. I hope you all enjoyed the talk.

And definitely if you are looking for some more insights, follow your boy. I can easily be found on this thing called the internet. And definitely LinkedIn is a great way to get connected here.

All right. So one of the first questions we have here is let me see for for someone who is getting into SR and the aspects of distributed systems can you tell me a general overview of what you would do if you have to start again? How would you get experience?

That’s a really great question. So, one thing that I would do for S sur first of all is building up two proficiencies and two skills. If you didn’t notice, I started off saying I’m a systems engineer or actually I did say I was a software engineer first, which I usually don’t, but I started off saying I was a software engineer, but I also am a systems engineer, right?

And I’m having two different skill sets in terms of proficiency. It could be a database administrator and also systems administrator, right? It could be a software engineer that manages front end but also backend that fills that skill set around site reliability engineer.

And what I mean by building proficiency is building a a level of proficiency to understand what’s going on in both lanes so that you can be prepared to provide that reliability support. The second thing that I would do is I would no doubtedly like be reading and reading even more that’s online about site reliability engineer. Let me see if I actually I got the book right here.

Talking about the book. Great resource. I’ll put it right here.

This no plug. Nobody’s getting no money off of this, but definitely it looks like somebody asked for what is a book recommendation, so I already dropped it out there. It’s called Site Reliable Engineer: How Google Runs Production Systems.

It’s a great great book, and I believe they actually have two series. Another good one, and actually y’all got me, you know, pulling out from my bookcase. If you’re learning about Unix and Linux administration, this is a great great great handbook.

And I can put these out somewhere afterwards. And then the last one I’m going to show is cracking the code. If you haven’t seen this one or read this one, this is a really good resource and definitely give you some insights into cracking the interview code.

But those would be the areas that I would be developing. And if you kind of seen the different type of books I have here, I have software engineering books, I have systems engineering books, and I have books, you know, on all type of topics. So definitely want to encourage you to check out some books that that’s where I would start off with.

And then the last thing is, hey, follow your boy, follow any newsletter or any topics around site reliability engineering. And definitely, oh, there’s a really good blog post or newsletter. It’s called Bite Bitego.

I don’t know if y’all, somebody can throw it in the chat room. I can’t pull it up real quick, but the person that created that blog post is the same person that created the system design book, and he used to work at Twitter, too. So, he’s a great resource, great technologist.

And definitely had some really in good insights into you know site reliability and designing systems at scale. All right, so I answered that question. Let’s see what else we got here.

Another question is I’m I’m trying to build a career at the intersection of AI systems and sustainability. What technical skills or infrastructure experience should I be prioritizing first to make myself a strong candidate for entrylevel roles? All right.

So, what I would say is the cool thing about where we’re at right now, especially with AI, is we are only and I and I have to, you know, go check the calendar, I believe, two or three years from when OpenAI released the first ChatGPT that everybody has access to. So, we are really at the beginning of this journey. And what skill I would encourage you develop is definitely utilizing this tool called AI to be more proficient, right?

And it’s not like, hey, AI is going to take my job. It’s going to be, hey, the person that develops the skills to utilize AI most efficient is going to take that job, right? And the cool thing about AI, especially in closing the gap on developing infrastructure experiences, AI can be used as a way to teach you infrastructure skills.

So much so that I have a mentee of mine that’s going through a program right now that’s only using AI to learn how the cloud works from first principle up to, you know, scheduling workload and some Lambda function, right? And really stretching that scope. Also on the infrastructure side, especially if you’re thinking about cloud, is certifications.

I know that certifications, at least when I was getting in the industry, early 2000s, was the hottest thing on the street, right? It really gave companies and insights into your skill sets. And I think this is coming back around.

One of the reasons why I decided to take the Red Hat certified engineer certification, it was the hardest certification at that time. I got certified in that in ’ 06. But once I got certified and Red Hat certified engineer doors started swinging open right is utilize certification as a way to you know get your badge and get your stripes on your shoulder so when you are engaging with companies they at least know that you have the capability and capacity to talk they’ll talk and if the certification is a real difficult one like the RACE they respect you when you first walk in the door so I don’t know if you’ve ever heard about that certification check it out it’s RHCE it’s red hat certified engineer u would be a way that I would encourage you all to level up.

All right. So, next question I have, is there a way that I can get in touch with you and contact you? I mean, LinkedIn, I’m I’m always online.

The internet is is, I’m I’m friendly on the internet, so definitely feel free to reach out. Just know that obviously with EES, there’s going to be a lot of people hitting up my LinkedIn. Be patient.

I definitely will reply back and get connected. U, but hey, take a take a look at some of the the cool articles that we have out there in our blog post. We talk about tech, we talk about AI, take a look at some of our YouTube content.

Every YouTube that I’m featured in, I do a tech talk. So, at the beginning of the podcast, I talk about something in AI and tech. And then I also travel the country and interview, you know, some technologists that I’ve had a chance to work with.

So, it’s easy to find me on these streets and definitely LinkedIn. Let’s get connected. All right.

Next question I have and might be the last one I have here is, I’m trying to become an AI data engineer. I’m I’m assuming I need to know everything from the hardware to making the model themselves. However, that seems like a lot for a single career.

Any advice? All right. So, AI data engineer.

So, anything with AI nowadays and don’t feel like you’re being called out specifically because it’s a AI data engineer. I think any role that has AI, ML, has you know data sciences has some implications that you’re going to be interacting with a model. There is a higher requirement or higher boundary to get into the familiarity of how the system works.

And the the the challenge is is that the bar right now is set extremely high. I mean, I don’t know if y’all noticed, but they were writing hundred million dollar checks for researchers from one lab from OpenAI to go to Meta, right? The the the bar is set really high.

And when that happens, you always find the need for you to be able to develop even more insights into the the technology to get recognition andor to get into the industry. So I don’t want to like have you hold back from this idea that you’re going to have to learn and this is about everything in engineering. We always have to be learning.

We whatever I look worked on five years ago is already deprecated and I got to learn again is get to the groove of really diving down into that niche and really sharpening the skills from beginning to end from yes understanding how hardware works and the GPU works compared to CPUs understanding what libraries you need to use to access the GPUs that let’s say a provider like Nvidia has from all right now I need to write a data pipeline using PyTorch or TensorFlow which are Python libraries to create a pipeline to train the model. Actually, I need to be able to access a SQL database to query this data set that we’re actually going to be training the model with. And actually, it’s written in SQL on my SQL and then it’s also stored in this thing called a map reduce, right?

Where it’s just some unorganized key value store like yes that is the spectrum from where the skill ranges from especially when you say AI data and engineering I do want to p encourage you to pace yourself continue to follow along find a community that’s really focused on AI data engineering so you have that community you I guarantee you can find it here at at the EES put it out there let other people know that you’re interested in and then Yes, I want you to go through that cycle to understand how you actually build a model so that you when you’re in an interview you can really dive into the topic and that that’s where passion really comes in and and actually kind of you know overlaps with the keynote that I heard today is you know focus and passion is key when when you’re really engaging with organizations nowadays because there’s just a lot of people on these streets. All right, I think we are at the end of our conversation. Anthony hopefully I did well and hold it down my guy.

No, Bobby, you did amazing. We we knew that already. Again, Bobby, we thank you so much for always just being a champion of CodePath and of our students.

We really really appreciate it. If you all can show Bobby just your appreciation in the chat, we thoroughly appreciate it. So again, we’ve come we’ve came to the end of this actual discussion.

Bobby, if you had one piece of advice, just one piece of advice, what would that piece of advice be for students that are afraid of this new wave of AI, you know, because it’s, you know, the talk is it’s going to, you know, consume all of the jobs. So, all of our recent grads and or our upcoming grads, they’re worried about this. Like, what is one piece of advice you will you will give them?

Yeah. So, ju just know that this iteration that we’re on, is something that I’m more excited than ever to see what will be created. You just have to keep the momentum.

Stay on this hustle. And what I mean by hustle is I’m not just applying for jobs through LinkedIn or through some job. I’m reaching directly out to companies and knocking on doors if I have to to get access.

You’ve got to stay on the hustle. And then the last thing that I will say is keep your ears to these streets because AI is changing so fast that within a week you can get lost and if you’re not staying up on this game you’re going to get left behind. That includes learning that developing that that includes like spending the weekend instead of going hanging out with your friends.

Actually, I’m going to go build an AI pipeline because I’m trying to build a like that type of commitment is what you’re going to need to stay up in the game nowadays. And I’m I’m only more encouraged that you all heard it from me and definitely here at CodePath as a supplemental education. And yeah, I’ll let your boy I got you for sure.

Bobby, I have one more question for you then I’m going let you go. Damn, Anthony. How you gonna get two questions, bro?

So So would you say is is is hyper important to ensure that they have a portfolio of some sort and or like their their GitHub is up to date. We’re able to see what they’re working on. Even their pain points.

Is that important for employers to to see? Yeah, 100%. And thank you so much for that, Anthony.

And this is the the career coach coming out of me now. One of the things that you all may not know is I’ve interviewed hundreds of candidates. I said I worked at Twitter.

You can imagine how many candidates I’ve interviewed. The first place that I go is I look on their resume for their GitHub link. I go to their GitHub and I pull up that page.

Do you know what I’m looking for? I’m looking for the color green. If I scroll down to the bottom, there’s a chart that tells me how active you’ve been.

If you’ve been active, all right. Yeah. Yeah.

Yeah. We can holler. If you haven’t been active, now I got to go figure out what you’ve been doing because you trying to apply for this job and we trying to pay you $200,000.

Like that’s where your activities, especially on GitHub, really is something that we honor like technologists like myself. And it’s not that we’re looking to go dive and get to what exact No, we want to see your progression. Any work that you’re doing is everything is work in progress.

Don’t wait until it’s done for you to upload it to. No, we want to see what it was six months ago to what it is right now. And that’s all visible on GitHub and those type of portfolio projects are key for you to get recognition nowadays.

So, if I can encourage you, make sure that chart down there as much as possible is green because then I know you really on these streets balling. Absolutely. Okay, Bobby, one more.

Just in two words. In two words, what is your biggest piece of advice for emerging engineers? If you can do it in two to three words.

So, there’s this there’s a saying in K that says now this word translates into I’m hungry, but it’s a different one. Is I want you to stay hungry. I want you to be gungoo.

I want you to be more motivated than ever because that hunger, that’s what’s going to drive and get you across this line because at the end of the day, there’s going to be a lot of people at this table, but the ones that hungry the most, the ones that really want to get it or the ones that are going to go are going to be the ones that come out on top. Absolutely. There you have it, Bobby D.

We appreciate you so much again. Again, and we’ll see you around. So, from there, we’re going to close out with a quick video from Bobby D.

And guys, make sure you take the survey. It just popped up. Make sure you take the survey.

And we’ll close it out with a video from Bobby D. That a great session. Oh my gosh, these breakout sessions are fire.

And I want to encourage you to attend as many as you can. But let me give you some information before we continue. So, please take a moment to complete a brief survey to share your feedback.

Remember, surveys are what’s going to make us get better, but also let us know what we’re doing really well. To see what’s coming up next, you can head back to the main stage and check out the schedule. See you soon in the next session.