Life Is But A Stream

Ep 24 - Boring by Design: Character.AI runs 5GBps on WarpSteam and saves 90%

Episode Summary

Maor Bril, Chaos Catalyst & Technical Staff Member at Character.AI, shares how a lean team built real-time safety and experimentation pipelines on WarpStream by Confluent—cutting infrastructure costs by 85 to 90% while running 5GB/s through Kafka-compliant, BYOC infrastructure.

Episode Notes

What does it take to run real-time safety checks on every prompt sent to an AI companion platform serving 60 million monthly users? For Character.AI, the answer was rethinking the entire data pipeline—without rethinking the budget. By building on WarpStream's Kafka-compatible, bring-your-own-cloud architecture, the team found a way to process massive prompt and response payloads in near real time, at a fraction of the cost of traditional Kafka.


 

In this episode, Maor Bril, Chaos Catalyst & Technical Staff Member at Character.AI, joins Joseph Morais to break down how WarpStream's agent-based, object-storage-backed design let a small team capture full prompts and responses for safety classification and model A/B testing, route client and server events without added infrastructure, and even build a custom feature store with Bento-powered Managed Pipelines—all while running on otherwise idle GPU-node CPUs.

You'll Learn:

Chapters:

[00:11] Character.AI Intro 
[02:04] Data Streaming Goodness 
[18:22] Beyond the Stream
[26:14] Quick Bytes
[30:06] Joseph’s Top Takeaways

About the Guest: 

Maor Bril is Chaos Catalyst & Technical Staff Member at Character.AI, where he has spent over two years leading infrastructure and product efforts beyond core one-on-one chat, including data pipelines, labs initiatives, and new product lines. He originally built much of Character.AI's data infrastructure and continues to work across platform, safety, and experimentation systems supporting the company's AI entertainment platform.

Guest Highlight:

"WarpStream is the most boring product I've ever got to work with. In infra, that's the biggest compliment you can ever give someone. It just works."

Dive Deeper into Data Streaming:

Get Connected:

Our Sponsor:  

Your data shouldn’t be a problem to manage. It should be your superpower. The Confluent data streaming platform transforms organizations with trustworthy, real-time data that seamlessly spans your entire environment and powers innovation across every use case. Create smarter, deploy faster, and maximize efficiency with a true data streaming platform from the pioneers in data streaming. Learn more at confluent.io.

Episode Transcription

[00:00:00] Maor: WarpStream is the most boring product I've ever got to work with. In fraud, that's the biggest compliment you can ever give someone. It just works

[00:00:11] Joseph: Welcome to Life is But a Stream. Today we're talking to Maor from Character.AI about how they were able to reduce their costs from 85 to 90% using WarpStream to build their safety and experimentation pipelines without adding any additional overhead. You want to find out how? Well, you've got to tune in. I'm your host, as always, Joseph Morais. [00:00:30] Let's get into it

Thanks for coming on the show, Maor. Let's jump right into it. Tell me about yourself and what Character AI does. 

[00:00:39] Maor: Sure. I'm, uh, Maor. I've been with Character for a little bit over two years, uh, where I primarily lead, um, e- everything that is not just, uh, one-on-one chat, which means that part of our infrastructure, a bunch of our labs product [00:01:00] stories as well.

But, um, uh, originally I started very he- heavy into the infant platform and built, um, a bunch of our data pipelines, which w- we'll discuss. Or as a Character.AI, leading AI entertainment, uh, platform with, with about, um, 60 million, um, monthly, uh, active users worldwide. Our platform usually empowers users to, you know, connect, learn, and, and, and tell stories, right?

Uh, some of them are using one-on-one chat, some of them are using, uh, some of our Imagine, which are multi, um, [00:01:30] uh, model offering. And, and basically w- we're making e- every user as an active participant in their own entertainment, right? So it, it's very, very personalized and, and very customized. That's in very high l- level, but basically our, our, our focus in the past, uh, three years, you know, where e- every other lab went towards, uh, uh, productivity, we went towards entertainment and making sure that the conversation is funny, right?

Um, so that gave us a very, very unique insight in how [00:02:00] users are, are engaging with AI when they're not just trying to solve problems.

[00:02:10] Joseph: We set the stage, so let's dig deeper into the heart of your data streaming journey in our first segment. So my first question to you is, what have you built or currently building with real-time data streaming? And let me just... Before you answer that, Character AI is a WarpStream customer, correct?

[00:02:25] Maor: Correct.

[00:02:26] Joseph: Okay. So let me, for the audience, so all the data streaming [00:02:30] enthusiasts out there, WarpStream, uh, is a technology by Confluent. We acquired WarpStream, to be honest. It, it's this fascinating technology. It uses agents that talk the Kafka API and then write directly to object storage, and there's a bunch of advantages to that.

Certainly can help with cost at massive scale. Uh, it also helps with some, you know, resiliency challenges, uh, and even some, um, you know, RPO zero type scenarios. But I don't wanna g- dive too far into that. Let me come back. Tell me what you've built or are [00:03:00] currently building on top of WarpStream.

[00:03:01] Maor: Most of our, our systems are built on top of one of the hyperscalers.

And because we serve a huge amount of u- users that are all concurrent, but, but also they're chatting a lot, um, our first biggest use case was that how do we get not just the messages, right? The, the actual messages are easy, but the messages plus the entire prompt that went into generating that response, because we need to collect all of that in real time.

We can run some classifiers in terms of trust and safety. Uh, [00:03:30] we need to, you know... I- if, if we collect the, the whole thing, we can use the shape to kind of A/B test different prompts and different models. Like, for example, with this input, right, how would this model, uh, would have behaved versus that one? A- all of these required a very massive scale.

Even though w- we did have an, you know, a solution that would, would scale very, very nicely fr- from our, uh, provider, but it was on the more expensive side, uh, e- especially at, at the scale we were, um, uh, looking into it. And also it would, would, would have required us to kind of [00:04:00] custom tail a, a lot of our existing infrastructure to work with that.

Bunch of research, w- I came to the conclusion that probably we need something that is Kafka compliant. So our, our, our go-to was, okay, can we get a, you know, a Kafka clusters, uh, running or other competitors? Until we, we ran into the WarpStream team, which, you know, sound... At, at first it sounded too good to be true, right?

Yeah. You know, it, it looks like Kafka. It talks like Kafka, but it doesn't have all the pain points of [00:04:30] Kafka. Um, you know, managing a big Kafka cluster is always hard. Originally, in my career, when I used to manage bigger Kafka clusters, like over 10 years ago It was probably a lot harder than, than it is now because tools evolved.

But simple things like i- in- decreasing the TTL of a topic or, um, uh, adding partitions or, uh, rebalancing the topic and adding more compute node, more brokers to handle that, the increased w- workload or, or s- or sometimes, you know, your, your requirements change and that [00:05:00] cluster can go down. But all the things, you know, on the one hand from, from the compute s- side of things became very easy as we, we go into Kubernetes, and us being a s- a smaller startup, we needed to focus on things that, that really mattered to us, you know, um, to quote acquired, like, um, did that make our beer taste better?

And WarpStream fit in into that, you know, solution perfectly. On the one hand, it looks like Kafka, it behaves like Kafka. On the other hand, it's, it's running as agents, as stateless agents on Kubernetes. Uh, we, we run them in, in an HPA, so they scale up and [00:05:30] scale down as the, the compute, uh, requirements go up and down.

Initial set of challenges, wh- which you kinda have to ch- change your mindset. Sure. But, and, but the, the, the huge other factor that it's running on top of S3 compatible storage- Right ... which means that you get all the benefits of S3 being globally replicated. If you wanna do a multi-region cluster, that comes easy with WarpStream and at a fraction of a cost.

What, what really made the thing, you know, fantastic is the fact that they [00:06:00] have the feature, um, I think it's called, um, Managed Pipelines, which is basically, um, um, Bento, uh- Yeah ... running em- embedded in the agent. So we, we got to build all of our data processing pipelines inside of the same, you know, um, uh, the same system.

So in, for, for our bigger clusters, we did separate those out in different agent groups just to kind of, uh, because the, the, the, the compute requirements were different between the brokers and, and, and the Bento agents. For, for the smaller workloads, it, it's all one big happy family of, [00:06:30] of agents that are going up and down, up and down.

WarpStream's the most boring product I've ever got to work with. In infra, that's the biggest compliment you can ever give someone. It just works.

[00:06:40] Joseph: We had, it's so, this is so funny. We had an engineer from Cursor on. Okay. Uh, and they're also a WarpStream customer. And his quote was, "Data goes in, data goes out. It just works."

[00:06:50] Maor: Yeah, it just works. 

[00:06:50] Joseph: And like, you basically just echoed that. And, and again, for, for people following along, if you're fans of Confluent, you know, we have this product called Confluent Cloud. It runs across all the [00:07:00] hyperscalers, and it's fully managed Kafka. It differs in certain ways, right? It, it, most of the clusters don't use object storage.

We do have one called Freight that does work in a similar way, and it, it's even less hands-free or more hands-free, excuse me, than WarpStream. However, WarpStream has a unique value prop, and that it is it runs on BYOC. So as Meir mentioned, the, the agents are running inside of, uh, Character AI's VPC, whereas, uh, with, with Confluent Cloud, it's all running inside of our VA- uh, VPC, like most SaaS [00:07:30] solutions are.

Um, so not only, you know, do you have this really kind of cool construct of the agents, you just worry about agents, you don't have to worry about anything else, because you're offloading really the complicated part is the storage, and the hyperscaler's taking on all of that. And as you mentioned, you know, those, was it nine nines or 10 nines of durability, uh, all the o- you get all of that.

It's all mul- it's inherently multi AZ. It's all the things you would want to... If you had data that you do not want to lose.

[00:07:53] Maor: Especially a- a- as, as an AI lab, when we buy, when we lease the H100 or, or what- whatever, like [00:08:00] our GPU nodes, they come with a lot of CPUs.

[00:08:03] Joseph: Yes.

[00:08:03] Maor: And, and, and usually it's more CPUs than we can use to actually serve inference.

So those CPUs are effectively free CPU compute that our WarpStream agents run on. Again, because we're paying for them any which way-

[00:08:16] Joseph: Yeah ...

[00:08:16] Maor: we, we regardless, then, you know, uh, for us it was like a win, win, win. A, we get to use a very cheap storage option, which is SCCC-compatible storage. We get to use our own BYOC, like our own compute nodes, [00:08:30] which we're paying for them, you know, as part of our commitment.

A- and then also there's a whole notion of the governance of data, right? It, it, it all remains within our own projects, right? Like no User data ever leaves our, our domain. That makes the whole story very, very com- you know, very easy to sell to, to the, the, our, our legal and, and, and, and security. It was like a win-win-win across the board.

And, and also, like, uh, uh, one thing I, I also wanna add is the partnership we had with, with the, the WarpStream team, where we needed some custom, [00:09:00] um, um, connectors. They, you know, they jumped on board and, and, and we, we, we partnered, uh, together and, and we were able to build a bigger pipeline, which is serving, uh, five gigs per second and, you know, just works.

[00:09:12] Joseph: Five gigs per second. Five gigs per second, that is so massive. Like- Yeah ... I think of five gigs per second just like in terms of the throughput. I'm a network engineer, or at least in a former life, and, man, that was fast, right? But-

[00:09:22] Maor: Yeah ...

[00:09:22] Joseph: that's five gigs of data. Processing data, uh, really incredible. So for all the, the AI companies [00:09:30] out here that are gonna watch this episode, you really should be using WarpStream 'cause, again, with Cursor, they had the same scenario where they were buying instances that were mostly GPU instances, and they only needed the GPUs from it, but they had all these unused CPUs, and they ended up just putting the agents there.

So I, I imagine this is a very common scenario for anyone that's working with LLMs at massive scale. It really seems like WarpStream is the way to go. So tell me more, what was the underlying data or technology challenge towards building real-time prompt [00:10:00] safety testing and experimentation? So I, I know you mentioned- Yeah

you know, you kinda already had Kafka going, but I'm curious if you took a step back, why was, uh, the Kafka API necessary for this?

[00:10:09] Maor: Because most of our, our, our, you know, our existing pipelines are, are using things like Spark and Trino and all things, all these things are native to Kafka. I mean, yes, y- you can run adapters to whatever custom solution, and yes, we used to run, like, very unique transformer job that we're taking from, from one data structure to another, from one engine to another.

Being [00:10:30] Kafka compliant really opened up, you know, so many options. Like, uh, you know, in a different use case where we brought in ClickHouse- for some of our a- analytics workloads so easy. You just say, "Oh, this is Kafka. I know how to work with Kafka. Done." It made... You know, because Kafka has been around for so long, it had a whole ecosystem grow around it, and, and being able to use Kafka without the, the pain points of, of, of running a, you know, traditional Kafka.

And again, it has, it has its use cases, right? So if you look at,

um, at the [00:11:00] core things w- we, we have to kind of, uh, really think about and, and reason about, is that you are giving away things in terms of the latency, of end-to-end latency, like the, uh, message going into the system on one end and coming out, uh, the other end. With, with, with traditional Kafka, you know, you get, like, very, very low, few millisecond latency end-to-end, whereas with WarpStream it can take seconds, especially if you pair it with, with Bento, where...

Or, uh, managed pipelines- Sure ... and, and you batch some of the outputs. Sometimes we'll [00:11:30] get, get to seven seconds end-to-end. Our use cases were perfectly fine with this trade-off, right? I mean, our use cases were, were not sensitive to latency. Uh, we needed to be near real time. We didn't have to be real time. So, like, especially for, like, other use cases where it was more Um, impression anal- analytics.

The, uh, use case where basically we, we're partnering w- with one of, um, the biggest providers of, of, of user-facing analytics. Uh, we have both a web app and a mobile app, and we wanted to measure so much more than we were measuring, but we're [00:12:00] limited- Sure ... by, by the amount of events per year. One initial solution we had is, is basically instead of sending directly to that, that vendor, we went through a proxy.

The proxy kind of separated the event types w- um, uh, we didn't need a- at that vendor, and everything else went to our data warehouse, again, through WarpStream. And it made it just adding more and more capabilities so easy. As you mentioned earlier, right, it runs in our compute cluster. So e- even routing inside the cluster is just [00:12:30] inside, you know, in Kubernetes cluster routing.

A, it's free. I mean, it's free, so you don't have to pay any egress b- because it, it's all routed w- within the same compute cluster, right? It, it, within the s- even, it can even be in the, in the same host if the agent is being, uh, being packed with, with, with your compute. So A, as I said, the network is blazing fast because it's all, like, inner cluster traffic.

Storage is cheap, and as you mentioned this earlier, the storage is backed up. We don't have to worry about bringing up the different brokers a- and, and, and all the other partitions, you know, [00:13:00] replicated, et cetera. So they're just there by default. And we have some use cases, uh, for the bigger pipeline with, with the prompts.

We actually, um, uh, stood up a Kubernetes cluster which had a bunch of, um, uh, more detailed, uh, eval and analytics w- that were running on a different compute. So different comp- but because S3 storage is global by default and multi-region by default, we get cross-cluster Kafka cluster or WarpStream cluster w- again, with no cost for ingress.

Yeah. And, and when we're doing five gigs per second, [00:13:30] egress can, can, you know, easily add up in cost. So as I said, the trade-off was A, the, the latency, but the added value of both the, uh, cost saving and sim- simplicity of, of operation, which for us was, like, the biggest win. It just made everything so much more worthwhile.

[00:13:47] Joseph: Right. To summarize, it was fast enough, massively scalable, easy to manage, highly durable, and kind of already fit into your architecture paradigm with having those free CPUs. I mean, that, that's a win across the board. I, I can't think of a [00:14:00] better scenario than that. I hinted at it in my last question. I know you alluded to it.

Why don't you take the audience through this real-time, uh, prompt safety testing experimentation? I know, you know, safety in, in terms of, uh, interacting with AIs is the top of, I think, a lot of people's mind. And I know you talked about using, uh, WarpStream to route these events to different models, but why don't you tell us about the full kind of scope about this real-time prompt safety testing?

[00:14:26] Maor: Whenever you talk to an LLM, right, um- [00:14:30] Except if you use an API, right? Your message is only a fraction of the actual data that's being sent to that LLM, because LLMs are, are stateless by, uh, by, by definition, right? Um, and every time you go through like, you know, your ChatGPT or Claude or, uh, whatever, right?

The actual engine behind it takes all the messages through, it takes the, um, uh, the system prompt, right, what- whatever other components-

[00:14:52] Joseph: Sure ...

[00:14:53] Maor: and builds the whole prompt that, that gets sent into the LLM, and the response gets, uh, streamed back to you. So for us, w- what, what [00:15:00] we needed to, to, you know, be able to reason about is, A, because the user only controls a certain small...

I mean, just a small part of this, right? We, we needed to be able to look at the whole picture- Mm-hmm ... right? Because we have a different system template, and if you're gonna use a different character, a different character definition into place, and we have some other features of improving memory, where we take summary.

So the, so the, the prompt is actually comprised of a lot of components where the user message, again, is only a small percentage. Um, uh, once we have this, you know, uh, bigger payload, A, we can debug [00:15:30] later if something went wrong. Like, for example, sometimes the, um, uh, models go a little off the rail, right?

You know, we get some gibberish or, uh, uh, whatever. So being, being able to look at the whole thing that kind of led to it or, or being able to replay it, replay it into, into a model in case of issues, kind of evaluate that, that itself is a very powerful tool. Being able to, you know, when, when we run, uh, when we're, we're trying to release a new, um, model, right?

Being able to re- rerun the exact same [00:16:00] input with, with the whole mess- like the whole thing into two versions of, or two models a- and being able to see, you know, um, a response A versus a, a response B, and being able to do that at scale. Uh, but also we use it to, you know, A, to, to run some safety, uh, uh, classifiers, as I mentioned earlier, right?

Uh, because sometimes the message itself is not necessarily, you know, it can be innocent enough, but in the context of the whole message, if you're in the context of the character, in the context [00:16:30] of, of everything that, that, that can allude to something that is uns- we have to look at the, the whole picture.

Um, uh, being able to do that at scale, you know, uh, is a re- really power, um, multiplier, especially when, when, when one of our core pillars of, of how we operate is make sure that AI is safe for our users. Absolutely. Um, um, so, so ha- being, being able to do that at our scale in close to real time is, we can only do this right now w- with the, you know, the help of, of, of, of the tooling, uh, we have.

[00:17:00] We can use, al- al- also use... And again, that, that's the cool thing about, about Kafka, that, you know, multiple consumers can consume at their own rate for- As, as they usually do,

[00:17:08] Joseph: yes ...

[00:17:09] Maor: for, uh, use case. So we can use the same pipeline, you know, as I, I mentioned, I already mentioned a few use cases, but we also use the same pipeline, you know, to understand what our users are talking about.

[00:17:18] Joseph: Right.

[00:17:18] Maor: Right? Like we- You know, we don't care about the individual user, we don't care about the message user, but we care about trends. We care about, you know, like for example, our specific fantasy theme, uh, or maybe [00:17:30] K-pop or what- whatever trending. And so, um, uh, we... That's a way for us to kind of in close to real time understand what, what is trending right now.

[00:17:38] Joseph: Sure.

[00:17:38] Maor: Right? Uh, as inferred, uh, by, by, by what our, our users are, are doing.

[00:17:43] Joseph: Next, we're gonna dive into how your partnership with WarpStream by Confluent solved your data challenges. But first, a quick word from our sponsor

[00:17:55] Announcer: Your data shouldn't be a problem to manage. It should be your superpower. [00:18:00] The Confluent data streaming platform transforms organizations with trustworthy real-time data that seamlessly spans your entire environment and powers innovation across every use case. Create smarter, deploy faster, and maximize efficiency with the true data streaming platform from the pioneers in data streaming

[00:18:26] Joseph: Now we'll go beyond the stream on why WarpStream by [00:18:30] Confluent was the right fit. Now we've established why data st- or data streaming was the answer. Let me... Why don't you take me through how you guys discovered WarpStream and maybe what that initial conversation was like. 'Cause I know, you know, anyone that's been around Kafka, you hear about WarpStream and you're like, "Eh."

You know, you already alluded to it. Yeah, exactly. 'Cause it's too good to be true. I, I'd really love to know, you know, uh, how r- WarpStream even got on your radar and maybe what was that first call like?

[00:18:56] Maor: As we said, right, we originally were looking at, we agreed that we needed a [00:19:00] Ka- Kafka based solution because it, it worked best and we're look- we're looking through the, the usual, um, I won't say usual suspects.

I'm not sure if that's the right terminology. Sure. But, um- You know

[00:19:09] Joseph: the usual players. Yeah.

[00:19:11] Maor: Yeah, yeah. The, uh, I mean, usual players and, you know, we were evaluating and w- we're having discussions and doing some cost estimations and, and we're realizing like, you know, how much it will cost us to go, you know, 10%, 20%, 50% the use and 100% in terms of, of the scope and, and scale.

This, and this was just, we only had like one big use case in mind. [00:19:30] One of our, my colleagues at the time- Mm-hmm ... you know, he went, "Yeah, I, I ran into this, uh, guy in, in New York. He told me about this startup called WarpStream and I really think we, we should goes out." At first my reaction was like, "Okay, cool.

You know, w- w- sure.

[00:19:43] Joseph: Yeah, we'll look at it."

[00:19:44] Maor: Yeah. "We'll look at it." Again, I looked at their, their site and it all looked like, you know, a classics startup tale. Like, e- e- everything is fantastic, everything is great. It looks so easy, so, so doing a proof of concept to, to pass on, right? I mean, like, you know- Sure

I, I, I'm, I, some, some of [00:20:00] these vendors to even, to even do a POC you have to spend a month in just kind of building the whole thing.

[00:20:04] Joseph: Right.

[00:20:05] Maor: With, with WarpStream, you know, you can run it on your laptop in like one command line and it, o- okay, functionally it works. That's awesome. Yeah. Let, let's run this, the, they gave us a Helm chart, I deployed it and we, we started pushing data through it.

And it worked. Yes, we knew, we, we, we realized what, what were the, the trade-offs. And yes, at first like I n- really have to have that 20, 20 millisecond latency and then like, um, why do I need [00:20:30] this?

[00:20:30] Joseph: Right. Question your assumptions, right?

[00:20:32] Maor: Exactly. And- Yeah ... and, and, and again, and, and e- every person, you know, I spoke with throughout my career when we're thinking about Kafka and streaming and latency, like the one thing you always hear like low latency.

And then, you know, when you think about if you're doing async, you don't really care about the latency. La- latency need to be good enough, but if you really need low latency, then do sync. If you're gonna go with async, then latency is secondary. And again, there are very good reasons why you wanna [00:21:00] do, you know, asynchr- or, or synchronous, but in many, many cases it's probably good enough.

And, and because the proof of concept went, went very, very well and we, we, we kind of scoped it out in terms of, you know, cost and whatnot, and it's a fraction of the cost of, you know, anyone else. Like it's about 10 to 15%, you know, total cost for a year. I was like, "Okay

[00:21:22] Joseph: That, that's a- Yeah ... that's like 85 to 90% reduction.

That's, that's wild.

[00:21:26] Maor: Correct. Correct. And, you know, they're like, "Okay, there's a, you [00:21:30] know, we'll try, we'll try them out. Worst case scenario, we wasted some time, but at least it will be fun." And, and you know, it, it wasn't all smooth sailing from, from the get-go as- Sure ... especially as we start, we started, uh, cranking up, cr- cranking up the data.

Getting the data out of, uh, of WarpStream at our scale wa- was harder, but it was harder not because WarpStream couldn't handle it. It's because the connectors we were trying to use at the time were not suitable for that scale.

[00:21:55] Joseph: You know, an 85 to 90% reduction in cost generally means you're gonna lose a lot of, like, [00:22:00] features.

You're gonna have to manage it a lot yourself, but you already talked about- Yep ... how easy it is. Correct. So imagine that, right? Yeah. How often in, in life do you get, especially in tech, do you get to reduce by that, I mean, even by 50% and you're not, you know, suddenly bringing on new engineers to handle that?

Very impressive. I had to call that out. Yeah, for sure. So I'm curious, was there any friction from your leadership or from the greater business about, you know, bringing on this new... At the time, I think WarpStream was r- was rel- relatively new. You [00:22:30] know, did anyone push back and say, "No, we need to use something that's more proven"?

I'm curious if you had any of the more personal political barriers and how you were able to overcome them.

[00:22:38] Maor: I think that, you know, e- e- eventually because we are a smaller team, then there's a, a, a bunch of trust that is kinda... If, if me or, or someone else, uh, you know, come up with, you know, uh, came to the table and said, "These are the choices, and we, we've looked into A, B, and C, and we chose to go with that one because of, you know- Uh, uh, reasons X and Y, [00:23:00] th- they're, they're very valid and data-driven, uh, uh, reasons, then w- we, we have the trust.

And we kind of cheated there a little bit because these solutions will cost you 10 times as much as this solution, so we might as well try that. So, uh, dollars and cents, yes. And, and, and dollars and cents, yes, but also, you know, it, it ticked all the boxes for, for, uh, security as well, right? Yeah. You know, we didn't have to bring in, like, a third-party vendor.

Sure. We didn't have to bring in new compute. It's all... We're using our existing compute, our existing [00:23:30] S3, um, compatible storage. It's using everything that is, you know, it was like a very easy s- solution. Like, the one thing w- we, we did is, I needed to confirm is that, you know, that b- because it's a BYOC solution, so the control plane remains on our providers and...

But, but again, it was a very easy thing to audit. But because it ticked all the boxes and was a lot cheaper, then, you know, it was worth a try.

[00:23:52] Joseph: That makes perfect sense. Now, I know you talked about that 85 to 90% reduction in cost, but were there any other final impacts of the [00:24:00] migration to WarpStream? Like, something that's outside of just the, the financial impact.

I mean, you already talked about another one, it was more secure by having it in VPC. Correct. But are there any other final impacts you can, that you can share with the audience?

[00:24:11] Maor: Um, I think that, you know, A, the fact that we're able to kind of build, build in a lot more use cases than what we had in mind because of the, the, the global nature of the, of, uh, of, of the actual storage.

So, um, uh, we're actually to, you know, a- able to, to take our, our, our, our training cluster and, a- and [00:24:30] use compute there that, that we, we, we use for the classifiers. And again, it all happened without having to kind of Figure out, oh, how, how get, how do I get the data from here to here? Do I go through our m- um, my data warehouse?

So I upload it here, download it there, et cetera. A- actually, another fun thing, we, we, uh, for a- another use case, we built our own feature store, and we actually, uh, wrote connector for Bento to be able to pull from our data warehouse to the, the new, uh, f- uh, feature store [00:25:00] that we wrote. And so that's running and, and it was, it was contributed back to, um, uh, to Managed Pipelines.

So we're actually a- able to use that to run both ETL and, and reverse ETL and, and so the whole solution where it really solved a lot of our ETL needs, which, which wasn't a use case. We, we know it was like an added bonus, oh, it can do that too. That's awesome. So it really solved a lot of those use cases that you don't even know that, that you're gonna have to solve at some point.

And the fact that it is so flexible, and I mentioned earlier, like, you know, for example, the fact that it's [00:25:30] in cluster, we needed a solution to emit, right? Um, I talked earlier about client events, right? That come from our, uh, app and web. We need a way to emit server event. One solution could be, okay, you create a gateway and every service needs to talk to the gateway.

Or everyone has a very thin library, which is basically a wrapper on top of a Kafka client that has some hard-coded, um, cluster addresses and just send the event there, ETL pipeline that knows h- how to handle the whole thing. So that made all these things a lot easier without having to, to stand [00:26:00] up different services and whatnot.

Just because it's running in our cluster, the net cost is zero, right? Because, you know, we pay for it any which way,

[00:26:08] Joseph: right?

[00:26:09] Maor: Um, uh, and it's secure and it's fast and, and, and it fits all of our use cases

[00:26:19] Joseph: Before we let you go, we're gonna do a lightning round. Byte-sized questions, byte-sized answers. That's B-Y-T-E. Like hot takes, but schema-backed and serialized. Are you ready?

[00:26:28] Maor: Yes.

[00:26:28] Joseph: What's something you [00:26:30] dislike about IT?

[00:26:31] Maor: Some things are harder than this, than, than they can be.

[00:26:34] Joseph: I think that's true about many things in life.

What is your hot take on the future of AI?

[00:26:38] Maor: Fascinating. I think agentic for the win. I think that the amount of unlock that is yet to come, fundamental. Like, I know that, and, and, and, and this is not a byte-sized answer, I apologize, but I come from also from a background of, of, of crypto, uh-

[00:26:52] Joseph: Sure ...

[00:26:52] Maor: for char- character, and crypto kind of took the moniker of Web3, but if you think about actual evolutions, that's, I mean, [00:27:00] AI, GenAI is Web 3.0.

The internet is different. How we as engineers work is different. Like, you know, if you took an engineers from 30 years ago and you put them to, uh, to be an engineer three years ago, maybe the languages changed, maybe the, the frameworks changed, but it's just, it was the same work, the same type of work.

Today, it's fundamentally different, and, and I think it will only go, you know... It, it, it enables us as, as builders to build, right? It, it, you know, all, all the things you said that, "Oh, I won't even start this to, [00:27:30] to test this idea because it'll take me, like, two or three months of, of scaffolding just to see something," now it's two hours and, and you can iterate fast, fast, fast, fast, fast.

So AI will bring it into a very personalized but also very hyper, you know, everything is, is cranked up to 11, so it's fantastic.

[00:27:45] Joseph: Where are you getting outside inspiration, whether it's from a book or maybe a thought leader?

[00:27:50] Maor: Um, I mainly read sci-fi and fantasy books, so-

[00:27:54] Joseph: Okay

[00:27:55] Maor: So, um, my, my, my recent, uh, f- uh, fab is, um, [00:28:00] Dungeon Crawl: Coral.

I, I, I don't know if, if you, you, you've read. If you haven't read, I recommend to listen. The, the, the audiobook is by far the best I've ever listened to.

[00:28:09] Joseph: Oh, wow.

[00:28:10] Maor: Yeah. Nice. I cannot recommend this enough.

[00:28:12] Joseph: Okay. Mm-hmm. I will check it out.

[00:28:14] Maor: Yes, uh, Jeff Hayes is amazing. But, but the thing is th- there are certain ideas there like, you know, uh, there's a whole system AI, so, um, and how every character there is going through their own personalized adventure.

That to me is like, you know, ca- ca- so it's, [00:28:30] there's like a bigger narrative, bigger arc, but it's so personalized, right? Everyone gets their own achievements and their own loot, and their own... And, and this is like when I'm looking at Character AI, yes, we have a very common platform for users to engage with characters, but we can make it so personalized that because, for example, an achievement that will crack me up because it was so funny to me, you know, you will say, "Okay, cool."

[00:28:55] Joseph: Um-

[00:28:57] Maor: Yeah, fair to you ... so, so, but, but, you know, but it's not, you know, like it's [00:29:00] all internal jokes and it's all private jokes and, but do it at scale. So that to me is like where I also draw a lot of ideas, you know, and, and being able to do it at scale, you know, we need the, the data for it, right? You know, otherwise it's just some person's idea.

[00:29:17] Joseph: Um-

[00:29:18] Maor: Right ... uh, so being able to, you know, build both things that, that the users want at scale. Awesome. Well, any final thoughts or anything to plug? Check out Character AI and check out, um, Kai FM, which is [00:29:30] my, um, my latest product under, uh, Ch- Character AI. WarpStream has been the most boring product I, I, I ever, ever got to use, and I, I say this with love.

Like, like really, it's, it's just been fantastic.

[00:29:42] Joseph: Yeah. I mean, uh, you know, I'm, I'm Italian American, right? And so we would say, with all due respect, WarpStream is the most boring product encountered. But I, we all know that's a compliment. Yeah. Um, thank you so much. Just this has been fascinating. I appreciate everything you've shared with our audience.

So [00:30:00] thank you again, Maor. And for the audience, stick around because after this, I'm giving you my top three takeaways in two minutes

Man, what a great conversation with Maor. Let's talk about my three top takeaways. The first one is kind of comes back to why they chose Kafka, right? So when they were looking at a solution for the data pipelines for their, you know, safety checks and experimentation, they realized that Kafka API [00:30:30] already plugged into everything they were using.

So it was a no-brainer to, to go with something that looked like Kafka because of the established ecosystem that Character AI already had. Now, they obviously shifted away from a more traditional Kafka scenario to WarpStream, and one of the reasons they did that, and my next takeaway, is they made... The WarpStream team made the POC as easy as possible.

In fact, you know, um, Maor talked about it, but you can run WarpStream locally on your laptop and test with it and, and, uh, actually get it up, uh, a lot faster than you could with many of the, uh, [00:31:00] containerized, uh, Kafka options out there. So I think especially for something that is so nascent as WarpStream, that was extremely important to get this, you know, hot startup like Character AI to even try the thing.

If you could try the thing in just a few minutes, why wouldn't you try the thing? But I think my favorite takeaway, and it kind of goes back to what we heard with our episode with Cursor, and, uh, with all due respect, WarpStream is boring, and that's a compliment. And why is it boring? 'Cause it just works.

Character AI was able to reduce their cost by 85 to [00:31:30] 90%. Their solution ended up more secure because everything was in their VPC, and they didn't have to give up any operational overhead. They didn't have to add more engineers or give up features, and they still had that massive reduction. Honestly, I can't think of a better compliment, and, uh, I'm gonna agree with Maor.

WarpStream is boring, and that's a good thing. That's it for this episode of Life Is But A Stream. Thanks again to Mayur for joining us, and thanks to you for tuning in. As [00:32:00] always, we're brought to you by Confluent. The Confluent data streaming platform is the data advantage every organization needs to innovate today and win tomorrow. Your unified platform to stream, connect, process, and govern your data starts at confluent.io.

If you'd like to connect, find me on LinkedIn, tell a friend or coworker about us, and subscribe to the show so you never miss an episode. We'll see you next time.