Life Is But A Stream

Ep 26 - Making Great Bets: How Extend Built a Serverless, Event-Driven Foundation with Confluent

Episode Summary

What does it take to run a commerce platform that's 100% serverless and fully event-driven? Jonathan Kropp, Director of AI & Architecture at Extend, shares how making early bets on Apache Kafka® and Confluent Cloud gave his team a durable, replayable central nervous system—and the freedom to focus engineering time on business value instead of infrastructure.

Episode Notes

Every architecture decision is a bet. Extend, the AI-native commerce platform unifying delivery, returns, exchanges, and warranty claims, made theirs early: 100% serverless, 100% infrastructure as code, and fully event-driven from day one. Those foundational bets now give the company something rare—the space to think about the future of its business instead of maintaining systems.

In this episode, Jonathan Kropp, Director of AI & Architecture at Extend, joins Joseph to trace the journey from cloud-native queues to Apache Kafka® on Confluent Cloud. He explains why the lack of persistence, schema enforcement, and replayability in traditional queuing services became a bottleneck, how event source mapping connects Confluent Cloud to their Lambda-based compute layer, and how a serverless stack lets Extend tie every revenue-generating event directly back to technology spend.

You'll Learn:

About the Guest: 

Jonathan Kropp is the Director of AI & Architecture at Extend, where he drives the technical vision and strategy behind a fully serverless, infrastructure-as-code, and test-automated platform. With over a decade of experience in software architecture, cloud computing, and event-driven systems, Jonathan specializes in designing scalable, high-performance solutions that empower developers and accelerate innovation. He leads a team of senior engineers and architects, fostering a culture of technical excellence and developer-first principles, and collaborates closely with executives to align technology with business objectives.

Guest Highlight:

"We've made a lot of great bets on our foundations, which has then given us that space to think about the future of our business, future of our platform, and the future of engineering. All of those great bets, Confluent being one of them, has allowed us the space to think about the future."

Chapters:

[00:52] Meet Extend: The Mission
[05:29] Data Streaming Goodness: Inside the Event-Driven Platform
[06:43] Serverless Strategy & Cost Visibility
[09:36] Why AWS Queues Hit Their Limits
[16:44] Beyond the Stream: Kafka as the Central Nervous System
[34:58] Quick Bytes
[37:56] Joseph’s Top 3 Takeaways

Dive Deeper into Data Streaming:

Links & Resources:

Our Sponsor:  

Your data shouldn’t be a problem to manage. It should be your superpower. The Confluent data streaming platform transforms organizations with trustworthy, real-time data that seamlessly spans your entire environment and powers innovation across every use case. Create smarter, deploy faster, and maximize efficiency with a true data streaming platform from the pioneers in data streaming. Learn more at confluent.io.

Episode Transcription

Editor’s note: This transcript is AI-generated.

Audio / Timestamps

[00:00:00] Jonathan: We're very happy with the decisions that we've made. We've made a lot of great bets on our foundations, which has then given us that space to think about the future of our business, future of our platform, and the future of engineering. All of those great bets, Confluent being one of them, has allowed us the space to think about the future.

All our engineers have to think about is producing an event and, like, executing it when an event comes in

[00:00:25] Joseph: Welcome to Life Is But A Stream. Today we're talking to Jon from Extend about [00:00:30] their journey to build a central nervous system, originally using native services from a cloud provider, but ultimately using Confluent Cloud to build a B2B2C systems around extending warranties and much more. I'm your host, Joseph Morais.

Let's get started

Thanks for coming on the show, Joe. Let's jump right into it. Tell me about yourself and what Extend does.

[00:00:57] Jonathan: Yeah. So I'm the director of architecture and AI [00:01:00] at Extend. Our team is focused on, like, how do we... Like, long-term technical strategy, innovation, and how we teach people to use tools and technologies, um, appropriately, um, so we're always making the right choice for our platform.

And Extend is an AI-native commerce platform that unifies delivery, returns, exchanges, and warranty claims inside of a single system, um, while utilizing Extend Shopper Intelligence to analyze customer behavior in real time, segment shoppers by value and risk, and dynamically [00:01:30] tailor policies accordingly.

[00:01:31] Joseph: So that's definitely... You've extended what Extend does. So we've- Yes ... known each other for a few years now, and I remember- Yeah ... the, the core competency of the company was the extended warranties, hence the name. Mm-hmm. But you just named a whole dif- a whole bunch of parts of the sales life cycle. So I guess that- Yeah

that has expanded over time.

[00:01:48] Jonathan: Yeah. So yeah, we started off as, like, AppleCare for all products everywhere. So how do we have a great experience for extended warranties or product protection, in our case? We saw a need there. There was a gap in the industry of [00:02:00] being able to move quickly and adapt to, you know, changing merchant needs, changing consumer behaviors, and how do we build a company that can adapt quickly to those changes?

That's where we initially started with product protection. Very successful there. Um, and then we moved into shipping protection as well. Lost s- damage and stolen goods. Um, again, successful there. And now we're jumping into, okay, how do we provide value to merchants and all those other touch points that they have with the consumer after the point of purchase?

Returns, exchanges, fraud and abuse. Um, and fraud and abuse is not [00:02:30] necessarily their, like, financial transactions, but, like, is someone taking advantage of your policies and causing a merchant to lose money? 'Cause, you know, returns are expensive. Number one, like, the shipping logistics all behind it, but also is it something they can actually resell once they've got it back?

They either take a cut on it, like a hit on that, or they can't do it and they have to destroy it in the field. Like, so there's a lot of different costs associated with all those other touch points that a merchant has to deal with after the point of sale. And so how do we allow them to focus on the things that actually generate [00:03:00] revenue for their business and help take the burden off of them for those other parts of the business that is more, like, in our wheelhouse because we have aggregate data across not just one merchant, but we can see signals across multiple different merchants.

[00:03:13] Joseph: Ultimately, you're delivering, like, excellent experiences and also, I guess, support of some of the most critical parts of the sales cycle, and all that is data-driven, correct?

[00:03:23] Jonathan: Yep. Yeah. Yeah, all of it is data-driven. Like, we, we try to pride ourselves on data-driven decisions. I

[00:03:28] Joseph: like that. [00:03:30] Um, so I think this is...

I ask this question in, in every episode I do, but I think this is, uh, extra apt because now with the extension of what Extend.com does, I think this answer has probably evolved. Can you describe a typical Extend customer?

[00:03:44] Jonathan: So what's interesting is we actually sell, we're like B2B2C. So we actually partner with merchants to sell on our behalf, or we integrate into their services to...

or into their platforms to provide complimentary services, like their claims processing, returns, exchanges, et cetera. [00:04:00] So a typical customer will actually go to a merchant partner that we work with, purchase a product, potentially purchase a protection plan with us. And after they have received pur- made their purchase, they can receive all their tracking, shipping updates, and file any issues with us.

And then if there's any issues with the product in itself, they can come through our claims experience and get, you know, made whole. And we can do that in a more streamlined, more efficient, more consumer-friendly way than a lot of, uh, other platforms out there.

[00:04:27] Joseph: Got you. So it's, it's retail businesses [00:04:30] buying your services, but then Extend...

I keep using the word Extend, and I don't even mean to do it. Um, extending those services to their customers. Yeah. You know, any, any person who's ever bought anything online probably has exposure to extend.com, but they would, they would be blissfully unaware of it.

[00:04:44] Jonathan: Yeah. I mean, so s- sometimes it is branded by us.

Like, you can see, like, Extend, like, powered by Extend, et cetera. Oh, cool. Sometimes, sometimes it is widely bought by the merchants themselves, and we just provide those services on the back

[00:04:55] Joseph: end. Cool. You meet them where they are. Awesome.

[00:04:56] Jonathan: Yeah.

[00:04:57] Joseph: All right. So we're gonna take-

[00:04:57] Jonathan: Again, it's, we're, we're trying to provide that, that [00:05:00] flexibility for merchants to plug into their systems and adapt to their needs and their consumer needs.

[00:05:04] Joseph: That's perfect. And I guess those needs, you know, as you've gotten more data and you see other opportunities to expend, uh, keep doing- ... extend your services to your customers, you're doing that. Okay. Very cool. Mm-hmm. Uh, it's great that we've gotten to know each other over the years, 'cause now I get to see the evolution of Extend.

So we set the stage. Let's dive deeper into the heart of your data streaming journey in our first segment. [00:05:30] So John, tell me, what have you and your team built or are currently building with real-time data streaming?

[00:05:35] Jonathan: So our entire platform is event-driven. Um, the nice thing actually is that we are 100% serverless, which means short-lived compute.

Like, we don't have any long-lived compute anywhere. Um, 100% infrastructure as code, so everything from the developer experience is, like, fully automated. 100% automated testing and top-tier developer experience. And then the next phase, obviously, like a lot of other companies, like how do we incorporate AI into the rest of our [00:06:00] platform?

But again, like this is all being data-driven Event-driven architecture is, is predicated on, like, having data in motion. So it just, it doesn't just sit still and you like, you request it when you need to. It's actually, the data itself is actually triggering other events going on throughout the system.

[00:06:14] Joseph: Yeah, 'cause I suppose you could be data-driven with batch, right?

But you wouldn't be as agile. So it's both- Yeah ... data-driven and as near real-time as possibly could be. So-

[00:06:24] Jonathan: Near real-time and flexible.

[00:06:26] Joseph: And

[00:06:26] Jonathan: flexible. 'Cause you know, it... Yeah.

[00:06:27] Joseph: Yeah. Um, so, you know, normally I don't [00:06:30] spend a ton of time with, uh, with customers about, you know, the, the how, right? Usually the show is really about the what and the why.

But because I know a little bit more about your architecture than I do most customers, um, I figure it's worthy, and you guys built something great. So if I'm not mistaken, your stack is more or less entirely serverless.

[00:06:48] Jonathan: A hun- yeah. Entirely serverless, and we've pretty much settled on, like, a single language or two languages basically to build our entire platform.

So it makes it, again, gives us a lot of flexibility across building different [00:07:00] abstractions and capabilities in our platform and allowing us to focus on delivering value as opposed to, like, maintaining and, and doing the operations of the system. 100% serverless, meaning that we get to... This is not necessarily saying that we have built everything in-house, but also relying on managed services through other providers to help provide that capability for us.

'Cause our, our goal is how do we focus on providing business value as opposed to maintaining systems?

[00:07:24] Joseph: Absolutely. You know, someone once told me, "The only thing worse than writing software that you could buy is, is [00:07:30] running a system that someone else will run for you." And you guys have really embraced that.

I think a lot of people from a, like, an idealistic standpoint, everyone wants to get to serverless. Everyone wants to use, uh, managed services to the best of their abilities. But you guys have really nailed it, and I know you... I remember you told me about, you can track the cost. What is it? Cost of an end-to-end transaction, something like that?

[00:07:50] Jonathan: Um, with- in the serverless, because everything is serverless- Yeah ... we can actually tie back all of our revenue, revenue-generating events back to our technology spend. So we can actually [00:08:00] see the cost that actually, the total cost for a revenue-generating event in our system and see that over time it's actually been decreasing.

Like, our cost per gener- for generating revenue is actually decreasing over time based on, like, our volume of traffic. 'Cause, like, whenever you have a serverless platform, everything you do costs money, and you can see that. So if you have, like, a huge spike, that's costing you money, and you wanna make sure you're tying that back to actual value.

And so anytime we can, we can see those spikes, we can actually go back and say, "Okay, we [00:08:30] need to adjust this." And that, that happens on a regular basis where we see something that we need to go adjust and we're able to respond to it very quickly.

[00:08:37] Joseph: To give an example, if some type of business action may extend $2, you could say, "That cost us 15 cents to deliver."

[00:08:45] Jonathan: Mm-hmm

[00:08:45] Joseph: Like something- From

[00:08:46] Jonathan: a technology standpoint, yeah.

[00:08:47] Joseph: Yeah. That's really cool. See, so for, for any of you tech leaders out there that are envious about that, you got to get on top of event-driven architecture and you got to get serverless. We talked about, uh, you know, what you've built, which is basically this, [00:09:00] um, underpinning of the substrate of event-driven architecture, all with serverless on top of it for all of your, your business logic.

Um, your engineers are able to focus on business logic 'cause you're down to two languages, which is wonderful. I can't think of how many companies have like, are all in on like four or five different languages and, uh, and every time I ask why, there's never a great answer. Um-

[00:09:21] Jonathan: Uh, I mean, that- that's a perk of being a younger company, too.

Again, we- we've- That is true ... we've made some great bets on like a solid foundation. The ear- the earlier you can make those decisions and, you know, [00:09:30] adapt to if they're the wrong decision, the better you're gonna be.

[00:09:33] Joseph: Let's go back to before all of this greatness was built. You know, what were the underlying data or technology challenges that initially prevented Extend from processing asynchronously across functions?

What prevented you, at least initially, from starting with a central nervous system, as I know you've described Confluent cloud, uh, and Confluent at, uh, Extend?

[00:09:52] Jonathan: So yeah, obviously we're, we're built on like AWS stack, so, you know, Lambda is our compute layer there. Um, we do have, you know, Dynamo and [00:10:00] S3 as like some like true data sources that we, that we utilize like heavily.

Communicating between different service boundaries, you know, SNS, SQS. Like we were starting to run into some of those, those issues where, you know, like the volume of traffic that we're going through was preventing us from actually... It was causing bottlenecks essentially. Not only bottlenecks, but then that data is not really persisted.

And so again, you fire off an event in SNS, SQS, and you have to... Your consumer has to be listening at that point in time for it to actually be able to process that information now [00:10:30] or in the future. Whereas we've... we saw the need that we would need to keep track of all those events that we're sending throughout our systems and be able to replay those in the future for different use cases.

Fraud signals, you're gonna come up with a lot of data points that you wanna use, and those are happening all in, in the past. Whereas the use cases and the different problems or the different ways you can go about solving those problems are gonna change in the future, so you wanna be able to use that historical data.

How do you actually have that data in a place that's reusable in [00:11:00] the future?

[00:11:00] Joseph: Absolutely. So you started, you know, with things like that are traditional queues, right? SQS, I mean. Mm-hmm. Queue is right in the name. You know, you guys started with this like we... Obviously we're gonna build our architecture around events, right?

'Cause SQS is great for events. But one of the things that was the sticking point is the lack of persistence in something like an SQS, and you couldn't do these like asynchronous consumption of events. Yeah. Because as soon as you consume something in a queue, it goes away. Is that m- in that, and I think you said some [00:11:30] scaling issues as well?

[00:11:31] Jonathan: Um, I mean, that's part of it, but then also schema enforcement. How, how do you actually do that at scale to make sure that you're not affecting your teams downstream? Like-

[00:11:39] Joseph: Got it. You know? So there were a couple of rough edges with an incumbent hyperscaler native service. Mm. And then ultimately you, you just figured out, okay, that's not gonna work.

So then take me... walk me through how you landed on Kafka, and then later on in the section we could talk more about how you guys landed on Confluent.

[00:11:58] Jonathan: Kafka, again, the, [00:12:00] the idea of Kafka is it's a persistent data store, like a persistent log file, essentially. Like, you can always wr- write to it, and then you can always replay it once it's been there, and you don't...

It's a, it's an immutable log, right? So that, that's one of those things that we were looking for is because we knew that data's future... is valuable in the future. And so that, that was one of the big pushes of, like, we need to get Kafka in here because we wanna have that replay ability, we wanna have that immutable log, um, because all those events that we wanna generate, valuable now, but, uh, we think they're actually even [00:12:30] more valuable in the future.

And so you, you can throw those into a data lake, et cetera, but, like, you know, ease of access to that data is a huge part of pro- solving the problem.

[00:12:37] Joseph: Absolutely. Like, if, if you think, "Hey, we're gonna need this data o- operationally," then it's a no-brainer to leave it in Kafka. Now, um, I'm curious-

[00:12:46] Jonathan: I, I think that's a great, that's a great call-out though, is like the difference between operational data and analytical data.

Yes.

[00:12:51] Joseph: And, and sometimes, you know, that, that data... I mean, obviously one is generally born off of the other. Mm-hmm. But they're not always utilized the same way, right? Having something in an aggregate table or something [00:13:00] that looks like a table, you know, some analytics system is not the same as, "Let me just replay this topic because I wanna train something," or, "I wanna understand where something was deficient," right?

And analytics- Mm-hmm ... really has a hard time. All that's been aggregated and all that, that operational data is just not there. So having that is, is extremely useful. And also, you know, I, I've seen scenarios where people pivot on operational sys- or, uh, sorry, analytical systems. Now, I think with open table formats it's a little bit easier now to just say, "Hey, I'm just gonna point a different engine at the same da- at the same data."

But sometimes that's not the [00:13:30] case, right? You, you end up going to a completely different system. And, you know, replaying from operational, from an operational state is one way to be able to shift your analytics, so.

[00:13:39] Jonathan: Not, not only just like changing, like, those, those analytical systems, but also whenever your service boundaries change.

You know, as a young company, you're still defining what your business is and those different service boundaries. And so as that expands and you have to start breaking things apart, how do you actually provide that data across those different systems without it becoming fra- uh, fragile?

[00:13:58] Joseph: Absolutely. So let's talk a bit [00:14:00] more about your service architecture.

So again, I, I nerd out on this. AWS Lambda is one of my favorite AWS services, if not my favorite. And there are multiple ways to connect Confluent Cloud with, uh, Lambda, um, the primary ones being either through the event source mapping or through the k- the con- the sync connector. I'm curious, which are you guys using today?

[00:14:18] Jonathan: Uh, we've been an early adopter of the event source mapping. So I think that, that's where we started off. Um, we saw a lot of benefits there. One of the biggest things that we like out of utilizing Lambda [00:14:30] is we get to focus on the execution and, like, what do we actually do with the data that comes in, and let AWS handle all the complexity behind the scenes of managing consumers.

Like, that's one of the biggest things with the event source mapper is that, like, Lambda then all we care about is, like, our log- business logic behind the scenes, and Lambda actually handles what a true consumer is, like the rebalancing, like- Mm-hmm ... um, the back pressure, all that kind of stuff.

[00:14:51] Joseph: Yeah, the

[00:14:51] Jonathan: scaling.

It's handled for us by our, our managed service.

[00:14:54] Joseph: Yeah, absolutely. I, I'm a big advocate of the event source mapping. You could do some filtering in it. Uh, as you mentioned- And correct ... [00:15:00] it, it'll scale out. Yeah. There's, there's a lot of ... I think in most cases it's the way to go. There's, there's-- I'm sure there's a couple corner cases where a sync connector, uh, maybe as like a, like a one-off or to fire, like, some type of mitigating Lambda or something.

In standard use case pattern, you know, event comes in, something needs to process it, event goes out, event source mapping is the way to go. I want to ensure that I've, I've summarized this correctly. Okay. So from the get-go, Extend solved value in building things inherently event-driven. [00:15:30] You started on AWS using some native services like SNS, SQS, realized that, you know, there were some rough edges in terms of the, the lack of retention, lack of schema enforcement, and then also, um, some scaling issues and things like that.

And then you realized quickly we have to move to Kafka, and then decided early, made the right bets in terms of Kafka as that central nervous system. You know, this is definitely the API we want to work with. Mm-hmm. Who we're getting our Kafka from is, is more of an open question. [00:16:00] Did I get that summarized correctly?

[00:16:01] Jonathan: Yeah. No, I think that's great. Great summary.

[00:16:03] Joseph: Next, we're gonna dive into how your partnership with Confluent solved your data challenges. But first, a quick word from our sponsor.

[00:16:14] Announcer: Your data shouldn't be a problem to manage. It should be your superpower. The Confluent data streaming platform transforms organizations with trustworthy real-time data that seamlessly spans your entire environment and powers [00:16:30] innovation across every use case. Create smarter, deploy faster, and maximize efficiency with the true data streaming platform from the pioneers in data streaming

[00:16:45] Joseph: Now we'll go beyond the stream on why Confluent was the right fit. All right, so we already knew that data streaming was ultimately the answer, or event-driven architecture. We like to say data streaming here, but they're kind of synonymous. Mm-hmm. And there's a reason why Kafka is the industry's standard [00:17:00] to address things like asynchronously, uh, routing functions, and for being a central nervous system.

But there's a lot to wrangle with Kafka, especially if you say we're running an open source. What made Extend choose to work with Confluent instead of, you know, building something of their own or going with some other service provider?

[00:17:18] Jonathan: So one of the things that we are really good at doing is we go and build things a lot internally to figure out what we actually want and the value that we want out of our partners.

So times- sometimes we'll go and build something and just so [00:17:30] we can dig in and figure out, okay, what functionality features do we actually care about? And our biggest thing with Kafka is we wanted to make this something where our engineers really don't have to think about anything except for how do I pr- like, I just wanna produce a message and I wanna, you know, consume it and, like, execute on top of it.

You know, obviously you got to think about, like, the schema that goes into the event a little bit as well. But the biggest thing was, like, how do we make this as great of a developer experience as possible for our engineers? We saw a lot of those capabilities at the time in, in Confluent, being able to, you know, automate...[00:18:00]

You know, this was a little bit of a struggle for us at first is like, how do we automate the infrastructure components for Confluent? Like, we actually had to do a lot of that ourselves, but over time that's become easier and easier to, to do. Um, also, you know, access controls. We don't want to have to necessarily maintain all those themselves.

If, if we can push that onto our providers and provide the right access controls based on our policies, again, we want to offload that work so we can focus on creating true business value and not focusing on operations or maintaining a system. And so we saw a lot of those benefits in choosing Confluent as our [00:18:30] Kafka provider.

[00:18:30] Joseph: So you, you mentioned earlier, like, the idea of building instead of buying, buying. Did you ever run open source Kafka at Extend?

[00:18:37] Jonathan: I mean, obviously like, you know, hosting it on Docker containers while we're testing them, doing a bunch of things. But again, we already knew ahead of time we didn't wanna host anything ourselves because we don't wanna manage that compute time.

Like, we don't wanna have to look at those systems and there's, and understand, oh, there's a problem here. We have to go make an update. We have to have downtime, et cetera. Like, we wanna offload as much of that operational and maintenance stuff onto our managed providers [00:19:00] so that we can focus, again, providing true business value.

[00:19:03] Joseph: Yeah. It's, it's interesting. You know, I've, I've talked to a couple founders of startups, Busi and Allium come to mind, um, that utilize like our startup program where we provide credits and things like that. Mm-hmm. And they have the same exact attitude. They were like, "We've, we've run... Like, we've been developers throughout our career.

We've, we've been part of on-call rotations. We're now starting our own business. We're not doing any of that."

[00:19:25] Jonathan: Yeah.

[00:19:25] Joseph: We're gonna build these managed services into our product and how we price and do everything for [00:19:30] all those same reasons. I know that, you know, y- you started the local development of Kafka.

You ultimately ended up on Confluent Cloud, and I do wanna talk more about that. Did you... Did Extend evaluate any of the other managed flavors of Kafka? I'm curious. I

[00:19:43] Jonathan: mean, we did evaluate MSK, um, the, the AWS hosting. Um, we've, we've looked at different Other either Kafka, Kafka providers or Kafka-like providers.

The value that we've seen in Confluent has kind of outweighed... Obviously, we, we look at a lot of different factors when comparing those competitors, and it's even n- easier [00:20:00] now to do with all the tooling available. So you just kind of pull that data in for you and create really rich dashboards to show you what the true cost of everything kind of looks like to make those, those decisions easier.

But again, each time we've done that, it's like, well, Confluent has shown that they're kind of ahead of the game in that sense, where, like, they're providing the right value where we, where we need it. If there's any gaps, having a great partnership actually helps figure out, like, okay, if we need something in the future, how do we actually get that done?

Like, there's work in the prog- I, I know there's stuff in the works at Confluent that is something that we're gonna take advantage [00:20:30] of whenever it gets released because it's something we've been wanting for a little bit, and we don't want to have to build it ourselves. Sure. So if you guys can pro- provide that solution for us, we're immediately gonna start adopting that.

[00:20:39] Joseph: Well, I love that you used the word partnership. I appreciate that. And I, I know you guys have been, uh, customers for Confluent Cloud as long as I've been here. Uh, and that's, uh, pretty early in the journey of Confluent's, you know, fully managed SaaS solution. Have you... Y- and you alluded to it, but are there any specific examples you can cite where, you know, the partnership really kind of stood out to you?

[00:21:00] Maybe you had some type of bug or, or feature need, and you were able to file that with a product team, and they were able to get to it, you know, expedite that, or perhaps, um, some guidance into migrations or onboard. I'm just curious if you had any of those truly partner-defining type of moments with your experience with Confluent.

I

[00:21:16] Jonathan: mean, our biggest challenge, I think, in the ecosystem of TypeScript was how do we actually... Like, the library that we're using. So there's a lot of libraries out there- Right ... for Kafka, but they're all kind of wrappers around, like, the C binaries. [00:21:30]

[00:21:30] Joseph: Yes.

[00:21:30] Jonathan: So how do, how do we actually utilize the Kafka libraries in a pure, like, TypeScript environment inside of a Lambda where you want to have, keep your packages as small as possible?

So that, that was a challenge of, like, working through figuring out Um, which packages to use. Working with the technical teams, like actually getting outside of like the sales teams and going and meeting with the engineering teams, um, the technical experts to understand, okay, why are you guys making these decisions?

What do we need to focus on to actually provide like the best experience for our engineers? And then how do we come to like the middle ground? 'Cause we [00:22:00] can't expect our partners to do everything for us, but we need to make sure that there's at least enough being done so that we don't have to do everything.

[00:22:06] Joseph: Did that scenario, did you guys get left in a happy place? 'Cause I... This, this is all coming back to me. I remember getting in this conversation, so I'm just curious if that all worked out.

[00:22:14] Jonathan: Um, I mean, yeah, it, it worked out in a sense where like we knew where you guys were going, like either like where you are or were not investing time, which then allowed us to make the right bet of like, okay, we need to go invest in this small area while like maintaining a specific [00:22:30] library for interacting with Kafka and Lambda.

But there's a lot of things you guys started building off from the role-based access controls, like from the infrastructure side of things that we could plug into, and we can rely on you guys for that. So there's two different aspects of it, like the infrastructure side and then the, the pure interaction side of Kafka.

Gotcha. Like producing and consuming. And so there was a lot more value of letting you guys own like all the infrastructure side of things and us plugging into more of those APIs that we didn't have to manage. And then there's a little bit of work that we have to do now, but also is made easier with the AI [00:23:00] tooling available to kind of maintain those, especially as there's new functionality being released inside our Confluent Cloud.

[00:23:06] Joseph: Absolutely ... it's

[00:23:06] Jonathan: not as big of a burden as it used to be.

[00:23:08] Joseph: I love that. I wish we could have just handed you exactly everything you needed but I'm glad that, um, the account team was able to give you the accessibility to product and engineering so that you guys could figure out what we were, were not going to invest in, and then you could then make those bets yourself.

It... 'Cause the worst thing you'd wanna do is, like, fund and start building something and then find out, like, two weeks later that Confluent's gonna release it anyway [00:23:30] 'cause that is the wrong way- Um ... to use your engineers' time. So I'm glad there was a middle ground.

[00:23:34] Jonathan: Unfortunate- fortunately, unfortunately, we tend to be about two to three months ahead of all of our providers in terms of, like, these are the things that we want.

We're gonna start building it to figure out what we want until they provide a solution, and we figure out either they're planning on building that or have... are getting ready to release that functionality. Um- Yeah.

[00:23:51] Joseph: I think it's just you guys make good bets, and, you know, people end up doing it anyway . But I think it's smart, um, to do that, like, to say, "Hey, we don't know if our v- vendors [00:24:00] are gonna do this.

Let's get ahead of it in case they don't, and then let's be nice and surprised when they ultimately build it ourselves, and then we can just sunset all this."

[00:24:07] Jonathan: Well, and then we can provide feedback on it. Like, "Hey, this is truly- Yes ... what we're looking for. We've already experimented with this." Like, um, we have the space to actually think about, like, the future of our business, the future of our platform, the futu- future of engineering.

It's a great bet for us to actually invest that time to figuring out what works for us and what doesn't, and provide that information to our vendors so that, again, we don't have to manage it going forward. We can provide that, that feedback and that [00:24:30] guidance to them of, like, this is what a technology-forward company is actually looking to utilize your, your system for.

[00:24:36] Joseph: I couldn't agree more. Um, yeah, I'm thinking about this, you know, with, with Extend being so ambitious and wanting to start day one with establishing a central nervous system and, and going all serverless, there's more than just tech to consider. I'm curious, did you have to get any buy-in from the business?

Like, was it initially all on board, like, this is how we're gonna build it? Was there any pushback to say, "No, we think a different architecture would be the right way to do it"? [00:25:00] Or in the vendor selection side of things, right? You guys started on AWS, and there's a lot of companies that decide, "Hey, we started with our hyperscaler.

We wanna go all in on hyperscaler services 'cause the economy of scale, 'cause of the APIs," et cetera, et cetera. I'm curious, was there any pushback either to go in to, to, to move forward to building, uh, everything out with Kafka, or was there any pushback, uh, adopting Confluent Cloud?

[00:25:21] Jonathan: Uh, I would say there was hesitation of adopting Kafka from, like, the engineering org, maybe because, like, they couldn't see the future value [00:25:30] of it.

It's like, "Okay, we already have an eventing system that works with SNS, SQS." Right. "Why do you want me... Why do you wanna enforce schema on me? Like, that just makes my job harder. Why do you wanna switch this platform and me think about things a little bit differently, 'cause, like, I already know how things work?"

So it took time to actually roll out and scale up Kafka and Confluent Cloud across the engineering org. Over time, people saw the stability that it provided. Uh, again, enforcing schemas earlier in the process And just making that, like, baked into our tooling has provided a lot [00:26:00] more stability across all of our other services downstream.

Like, we're not breaking as many services, as many contracts downstream and causing out- like, uh, not outages, but, like, degradative services of a schema change, or one team's thinking about one use case and not considering, like, use case, like, five or six that someone else is doing. Getting those foundational pieces in place has actually provided a lot more stability across the platform.

And the other side that is, has started to click for people is, like, actually having the replay capa- capabilities. Like, the use cases they were thinking about in the first, [00:26:30] uh, at first were just like, "Oh, I just need this real time. Like, I just need it now. I don't think, I don't care about the future." As we progressed, people have found it more and more useful.

Like, "Oh, now I see it's more u- it's more valuable. I need r- I need all that data. Let me replay it," or, "I had an issue here. What was the issue? I need to replay this, like, small set of data." Like, so they're seeing the value of replayability, so it just took time to run into those scenarios to show and realize that value.

[00:26:55] Joseph: Yeah. That, that makes sense. So it's a lot of the, you know, the classic, you know, write your unit test [00:27:00] first, right? If you're an engineer. And nobody writes the right test, test cases, but they'll end up saving your butt- Mm-hmm ... eventually. And it's, it's similar, you know, scenario. It's like, hey, schemas are hard.

Uh, they make things more difficult, but the benefit of them, especially when you think about evolution or downstream consumers or, you know, expanding or s- sc- uh, uh, creeping use cases and scenarios, uh, absolutely you want to start with schemas. I can't tell you how many times on this show someone said, "Start with schemas."

Like, sche- Yeah ... it's the most simple, like, [00:27:30] trivial thing, but start with schemas. And then the other one, the, the whole replayability, that's, like, a perfect example. It's like, "Hey, why do we need this? Like, we're, we don't need this data." Again, until you do, and then suddenly your life is so much better. So- Mm-hmm

it's great that you guys had the vision to push the company in that direction and to, you know, enforce things like schemas, even if, you know, the, uh, the people that had to, you know, enforce or, or, or be subject to those edicts, uh, ultimately didn't see the value initially. But then ul- eventually they [00:28:00] did, and then e- everything's copacetic and someone's like, "Well, Jon, we made the right decision.

Thank you

[00:28:05] Jonathan: for that." Yeah. E- e- every technology change comes back to you, you have to give everybody their aha moment. Like, like, what, how does this apply to my job, to my role, my function? How do I utilize this? And then they're like, "Oh, now I get it," and it's, the conversation is so much easier after that.

[00:28:19] Joseph: Yeah. Once, once you're addressing their pain, suddenly they care about it. Yes. I get that.

[00:28:24] Jonathan: As someone... I heard someone make a comment that, like, engineers are always willing to, you know, offer a new [00:28:30] solution to a problem, but they rarely wanna change the way they solve a problem.

[00:28:33] Joseph: That's-

[00:28:33] Jonathan: So they, they rarely want, they rarely wanna- Very true

take somebody else's recommendation.

[00:28:37] Joseph: That's right. Do it my way. All right. So we are years past the decision of implementing Confluent Cloud. You guys have built out this incredible serverless system. Um, you've expanded, uh, the functions of what Extend does over the years we've known each other. My question is around the final impact, right?

Like, and I think I've kind of cited it, but I want it in [00:29:00] your own words. You know, what would you say- Five years later, after making all these decisions and implementing these, you know, this, this stack, did you guys do the right thing? Are you happy where Extend is right now? And, uh, and I'm eventually gonna ask you about the future, but we'll save that for another question.

[00:29:14] Jonathan: I would say, yeah, we're, we're very happy with the decisions that we've-- Um, I think either we mentioned this earlier or, uh, another day we were talking, but we've made a lot of great bets on our foundations, which has then allowed us to-- given us that space to think about the future of our [00:29:30] business, future of our platform, and the future of engineering.

So, like, all of those great bets, Confluent being one of them, has allowed us the space to think about the future. Because again, it's one of those things that is just all our engineers have to think about is producing an event and cons- and, like, executing when an event comes in

[00:29:47] Joseph: I love that. Making, making great bets.

That is definitely-- We're definitely gonna incorporate that into the title of this episode. So, um, could you share some advice or lessons learned for leaders like yourself that are just starting to tackle data streaming or th- [00:30:00] or considering pivoting to data streaming and event-driven architecture?

[00:30:03] Jonathan: It's all about setting the foundation.

Set, set that foundation early. Just what we talked about, schemas first. Like, if you don't have schemas in place, like, you're gonna have a bad day. Setting up schemas and then offloading as much as you can to your managed providers, because what is your core business? Make sure that the time you're spending on solving problems is spent on what your core business is.

If you can offload some of that work onto, um, your managed providers, like, that's a win in every case. [00:30:30] The goal is, like, space to think about your business, not necessarily somebody else's technology.

[00:30:35] Joseph: Yeah.

[00:30:35] Jonathan: Yeah. Technology sh- technology should never be a barrier to getting, to providing business value.

[00:30:40] Joseph: Technology should never be a barrier to providing business value. I like that. Um, tech should never be the barrier. Uh, and I agree with you. You know, most of these things that you mentioned, whether it's like ephemeral compute like Lambda or event-driven architecture, it's undifferentiated heavy lifting, right?

It's like we, the providers, have this figured out. We have it [00:31:00] figured out at massive scale, at a scale that you're never gonna have to... I mean, I hope you get to the scale that, that, that all of our customer base or the entirety of our customer base looks like. Um, that's like a problem for us to handle and it's, that's our core competency.

That's our core business value, not yours. So I'm glad you're partnered with us. So you've already alluded to some of the AI things happening at Extend. Can you talk about those a little bit? I'm curious. And then maybe what the future and the vision for data streaming at Extend is into the future.

[00:31:27] Jonathan: Yeah.

So I mean, obviously like every [00:31:30] technology forward company right now, we're figuring out how do we incorporate AI into as many processes as possible. But the goal there is to provide effective use of AI. So I mean, we do have traditional ML models that we're still utilizing. It's not just all LLM or agentic based.

Like, that's still, there's still a lot of value in, in the ML space in providing that capability for your business. Um, from the agentic side of things, like, okay, how do you reduce or how do you provide better quality output by condensing all the dead space in between those interactions or death [00:32:00] by a thousand cuts that happen whenever you're trying to deliver something?

So like from an engineering perspective, like just like, like the SDLC process, for instance, like there's a lot of steps in between And then conversations are gonna go back and forth between, you know, product ideation, design, like user experience design, um, solution designing, and then going into, okay, how do we actually break this down into stories that are manageable?

How do we actually implement those stories? How do we deploy things? How do we test it? How do we monitor it? Like, each phase of that [00:32:30] SDLC process, there are things that you can utilize AI for to help, again, remove all that dead space in between those interactions. That's one area that we're, you know, focusing on.

Also, just think about all those mundane, boring tasks that you do. You know, there's a lot of hype around, like, coding and, like, all these, like, hot topic of where you can apply AI. There's so many other places where you can just apply it to, like, the boring side of your business- Sure ... and get efficiency gains out of.

Um, 'cause again, it's reducing all that dead space in between.

[00:32:58] Joseph: Yeah. I'll give you a good example [00:33:00] of that. We have a, we have a project, we built an agent that it puts, that tags metadata to, to documents. That's it, right ? Uh, but it does a really good job, and it's very hard to handle metadata across tens of thousands of documents.

But if, if you say, "Hey, I want all the documents that are related to, um, I don't know, uh, event-driven architecture," that's a pretty wide bucket, but you get it. And then you, then you had that metadata, now you suddenly, you're, you're down to 1,000 documents. Um, but little things like that. And I like that you mentioned kind of trying, you know, looking at the SDLC and looking at little things that you could take away from [00:33:30] it.

I think of, uh, when people want to, uh, migrate off a monolith to event-driven architecture, there's a strangler fig pattern where you kind of do, like, one little function at a time. You kind of, you know, pull one piece of the monolith until the monolith is gone. It kind of reminds me of that, where you're just, you're not like, "Let's not complete-replace the whole thing.

Let's look for things that are easily repeatable or that perfectly work with the ML, uh, like an ML paradigm or an LLM. Start to move some of those things over to make ourselves more efficient." I think that is a very [00:34:00] smart way to approach the future with agentic AI. Now, um-

[00:34:03] Jonathan: I mean, you, you even mentioned- Yeah

like tagging the metadata, like knowledge bases, right?

[00:34:07] Joseph: Yeah.

[00:34:07] Jonathan: So, like, a knowledge base is an extremely useful thing for a business, like, not just engineering, but also, like, the broader business as a whole. That's going back to, like, how do you find the right information? So choosing the right knowledge base solution- Yes

is a huge benefit to the company. I know we both use the same knowledge base, uh, solution, um, for, for the company wide. From, like, day one when we rolled that out, like, it's been a high [00:34:30] usage tool in our company, 'cause everyone saw the value of just, like, how do I get data in the right spot across all these different systems?

[00:34:36] Joseph: Right. How do I search across 100 different, uh, areas of, uh, of knowledge base? Uh, it's really, it is a huge challenge. Well, that's a fantastic vision that you guys had both initially at the onset of building this out at Extend, and then many years later. Um, very impressive stuff, John.

So before we let you go, we're gonna do a [00:35:00] lightning round. Bite-sized questions, bite-sized answers. That is B-Y-T-E. Like hot takes, but schema-backed, I know you're impressed by that, and serialized. Are you ready?

[00:35:08] Jonathan: Let's go for it.

[00:35:08] Joseph: What's something you dislike about IT?

[00:35:10] Jonathan: Anytime there's human intervention. Automate as much as possible.

[00:35:13] Joseph: I like that. No peop- Get, get my people out of my systems. Uh, what is your hot take on the future of AI?

[00:35:19] Jonathan: Uh, I think the biggest thing is the journey is part of the value. Like, it's not just about getting a single solution the first time you use it. Like, you have to fail, and you have to experiment to understand how it actually applies [00:35:30] to your business.

[00:35:30] Joseph: That's a really good answer. What's a non-tech activity or hobby that's impacted the way you think about data?

[00:35:35] Jonathan: When you're in a public space, like an elevator or just, like, a public space in general, don't look at your phone. Look at all the people around you. Like, what are they doing, and, like, how does that data, you know, apply to different businesses, and, like, is there a way to create, like, a digital twin that, you know, mimics human interactions in certain spaces?

I, I think that's something that I've had a couple of those conversations lately, and that's interesting to me.

[00:35:58] Joseph: That's interesting. So putting [00:36:00] analytical approach through mundane human scenarios. Okay, I'm gonna start doing that. Where are you getting outside inspiration from? Are, are you following... Are you reading any particular set of books or perhaps following a thought leader?

[00:36:10] Jonathan: Talking to as many people in, like, startup worlds and, and just, like, other leaders across other companies is a big thing that I have to do, especially, um, not being based in Silicon Valley, it's a lot harder to do that where I'm at. I think there's one book that I read rec- recently that, um, I thought was really great.

It's called, like, the re- it's called Reshuffle: Who Wins When AI... When AI Restacks the [00:36:30] Knowledge Economy. And it goes through a different... couple different use cases, like, the first one being kind of like container, shipping containers. Like, that was a technology shift. Who's actually the winner, uh, when you start thinking about, like, second- or third-order effects of that change, of standardizing on a shipping container?

What does it mean to adopt, like, ML fundamentals when you are a, a commerce pl- or a commerce company like Walmart or Kmart? Like, why did Walmart win and Kmart didn't succeed nearly as well as, as Walmart? It's like one of them [00:37:00] embraced technology and built a foundation on top of it to be more efficient, more agile, and more flexible.

So it goes through a lot of those examples to just think about, like, it's not just the surface layer. This is the w- problem I'm solving today, but what, what do you need to enable as a foundation to make that technology actually work for you?

[00:37:15] Joseph: Is technology something you use, or is your company immersed in technology?

Mm. And I think obviously no one Kmart didn't realize it. But if you're the latter, uh, you're probably gonna have a lot better outcomes. Um- Yeah ... especially in 2026. You [00:37:30] have any final thoughts or anything to plug, John?

[00:37:31] Jonathan: Um, no plugs. Um, I mean, my... I would just say never stop learning, never stop exploring. The world's always changing and, you know, there's always something to learn.

[00:37:41] Joseph: Yeah, and I don't think it's ever been more true than it is today. Thank you so much for joining me, John. And for the audience, stick around, because after this, I'm giving you my top three takeaways in two minutes.

That was a great conversation with Jon. Here are my top three takeaways. So the first [00:38:00] one is, uh, Extend's journey to ultimately, uh, standardizing on event-driven architecture, right? They started with traditional, uh, native services on their cloud provider like SNS and SQS. Uh, the problem is those didn't have the scale that they needed.

They didn't have the retention that Extend ultimately wanted and those durability of events, and they had no way to manage their schemas. And, uh of course, their engineers pushed back about, uh, why do I need schemas? You're making my life harder. Um, but as Jon said, you've gotta give your engineers that [00:38:30] aha moment in order to understand why they're, um, dealing with the pain upfront.

So it was wonderful that not only did they find event-driven architecture, but they immediately embraced it in a way that would, uh, set them up for success in the future. I like that Jon mentioned setting the foundation, right? Like with schemas and offloading as much as you can to managed providers. That allowed Extend to focus on their core business.

And ultimately, in every business we're in, unless we are a service provider, [00:39:00] we are providing value to the business. There's business logic. That's what we should be writing, the things that only we know about our business and our industry. And I'm gonna quote Jon, "Tech should never get in the way of business value."

And that's why pr- service providers like Confluent exist today. We run event-driven architecture at massive scale, so you don't have to, so you can focus on that business value. And the last thing I wanna point out is the partnership between Extend and Confluent. Extend had access to product at Confluent in [00:39:30] engineering in order to make the right bets about how they want to spend their engineering hours.

Should we build this thing, or is Confluent the provider gonna bring that out in a, in a few months? And ultimately, they met in the middle where Confluent was doing some of the, the, uh, development of the TypeScript library and, uh, Extend was doing the rest. Just an incredible conversation around why event-driven architecture is so important, the value of managed services, and what a fantastic partnership can look like.[00:40:00]

That's it for this episode of Life Is But A Stream. Thanks again to Jon for joining us, and thanks to you for tuning in. As always, we're brought to you by Confluent. The Confluent data streaming platform is the data advantage every organization needs to innovate today and win tomorrow. Your unified platform to stream, connect, process, and govern your data starts at Confluent.io.

If you'd like to connect, find me on LinkedIn, tell a friend or coworker about us, and subscribe to the show so you never miss an episode. We'll see you next [00:40:30] time