Why Smarter AI Agents Still Break With Juhi Parekh
In this episode of AI Explained, we are joined by Juhi Parekh, GM of Key Frontier AGI Accounts at Turing. Juhi brings experience across the full AI stack, from applied AI and foundation models to data infrastructure, with prior product roles at Apple, Amazon, Niantic, Spatial, and Samsung Research US, where she focused on commercializing frontier AI.
She explains how Frontier Labs curates hard datasets that maximize information gain rather than raw difficulty, why the sweet spot for reinforcement learning tasks is problems frontier models fail at least 30 percent of the time, and how long-horizon, real-world workflows are pushing agents to take on more complex work. She also shares the usual suspects when agents break in production (inaccurate tool calls, consistency gaps, permissioning, and output format), why training a capable model and building a reliable agent are two different problems, and why the winners will be the organizations that safely expand agent freedom as guardrails improve.
[00:00:00]
Introductions
[00:00:06] Buddy Brewer: Thank you for joining us on today's AI Explained on why do smarter AI agents still break? Buddy, VP of Product here at Fiddler, and I'll be your host today.So we have a really, really special guest on today's AI Explained, and that is Juhi Parekh, GM of Key Frontier AGI Accounts at Turing. So, uh, Turing's a fascinating company. You know, we think of the, the Frontier Labs as the absolute leaders in AI pushing the cutting edge of everything that's possible. But even those companies need to turn to other people for help. And when they need to turn to other people for help, one of the companies they turn in- turn to is Turing, uh, where Juhi is from. And so we're very fortunate to have her today to learn from her experiences working on from very many different angles. So welcome, Juhi. To kick us off, why don't you tell us about yourself and your work? How did you end up seeing both sides of frontier AI and what labs build and how those models perform once enterprises deploy them?
[00:01:10] Juhi Parekh: Hi, everyone. Uh, nice to meet you. I hope I'm audible
[00:01:17] Buddy Brewer: Hey, you sound great
[00:01:18] Juhi Parekh: Okay. Okay, great. Yeah. Um, yeah, so to tell you all a little bit about me, I'm a full stack AI generalist where I bring depth in AI, in the AI, AI ML stack across applied, uh, foundation models and now data infra at Turing. I have breadth across roles and scale because I've worked at startups and big tech like Apple AI ML, um, Samsung Research US, and Amazon prior to joining Turing.
[00:01:47] Um, by degree and profession, I'm a data scientist turned zero to one AI research product manager with an MBA. And over the last five years, uh, I've spent that commercializing frontier AI into revenue in some shape or form. Um, to answer your question specifically, I work at that seam between frontier model development and real-world deployment.
[00:02:12] So at Turing, I work with frontier labs on post-training, where we build hard datasets, benchmarks, and reinforcement learning environments that expose where the models fail. And then on the enterprise side, I see what happens when those models are connected to, you know, real tools, data, permissions, business processes, and just real-world workflows.
[00:02:36] Uh, before Turing, I led product roles at Apple, Amazon, Niantic, Spatial, and Samsung Research US, where I saw that stack, uh, across... Where at, at Niantic, for example, I was building, I was working on the foundation model, so I... That helped me understand what researchers actually care about and then shape that vision and strategy, which helps me.
[00:02:58] And across the five years, like, and across all the three, uh, horizons that I see, like applied foundation and, um, infrastructure, I keep running into the same two questions in this industry. Like one is: What does it take to build a good model? And typically, there are three things. One is data, second is researcher talent, and third is compute.
[00:03:23] And the second is: Where does that model actually solve a real-world problem? And at Turing, we in- work on both
[00:03:33] Buddy Brewer: That's really fascinating. That's really fascinating. You know, I, I talk to people who have done a lot of AI in the research space, have not been involved in building and commercializing AI in a commercial setting. and then I talk to people, and I'm one of these people, who builds applications on AI in a setting who don't have a background in AI research. And, uh, I think the space is moving so fast and, and, you know, as a, as an industry trend, it's been adopted by organizations large and small so fast that we're all racing to catch up in some way and on some dimension. It's really fascinating that you have a background in all of these different angles.
[00:04:14] How-- I'm, I'm curious that has helped you, if you have any observations on how understanding more of the fundamentals have made you better able to commercialize applications, or how understanding how the technologies get applied in a commercial setting have helped you help these frontier labs build better foundation models
[00:04:35] Juhi Parekh: Yeah, that's a very good question. Um I think
How Dual Fluency in Research & Products Sharpens AI Thinking
[00:04:39] Juhi Parekh: I'll break down my answer into two parts. Uh, there are certain skills that you need on either side. It can either be a very standard product management approach where, hey, there is a problem and there is a customer need, and to solve that problem, AI is a mechanism that we leverage.
[00:05:00] Like, you can solve that problem in multiple ways. AI is just one of those. And this... But what I see happening... And that was traditionally what was happening before this, like, whole ChatGPT, um, launch in November twenty twenty-one or two. One, I think. Yeah. Uh, and so, so that is traditionally how it was working before.
[00:05:26] But then after that wave, how it's been working is that, hey, we have this brilliant piece of research, and now let's bring it to market. Uh, so to, to answer your question specifically, there are, like, two muscles that you need as an individual and, and that's where the overlap, like, sort of... Or that, like, being able to zoom out, think of it from, like, an end-to-end, like, perspective, and then diving deep into, like, the specific vertical that I might be working on has, is, has been particularly helpful for me.
[00:05:58] So for example, like, at, uh, Apple, uh, where one of the things I was working on was, like, in fact, like, what are the use cases that would make augmented reality experiences mainstream from camera and search? Uh, so that was almo- I... That, that almost felt like a venture investing, uh, skill set, where you have to cut through ambiguity to find user wants that could become monetizable needs.
[00:06:29] It's like, hey, I'm looking through, like, eight different, like, verticals, could be education, could be healthcare, and then tying it to, like, the business strategy, and then identifying what might be specific use cases, and then building, like, engineering, like, prototypes around it, leveraging, like, the advanced research that we were developing.
[00:06:49] So that's the more, like, that's the pure play commercialization lens from a product point of view. And then commercialization in, like, uh, the infrastructure space, like for example, data infra, like, that looks very different,
[00:07:03] Buddy Brewer: Mm-hmm.
[00:07:05] Juhi Parekh: uh, uh, since you are not just deploying agents. When you deploy those agents, you identify, okay, where are the agents fa- failing in real world?
[00:07:12] And those can be training signals to actually, like, build data sets that can help improve the frontier models. Does that make sense?
[00:07:21] Buddy Brewer: Yeah. Yeah. No, it's fascinating. A- a- and like I said, I don't think that, um, it's so common that we run into people, uh, who have that perspective of ... And I always think that people are interesting who can approach, uh, uh, a large space with a perspective from different sides, right? So on the one hand, like the one we just talked about is from a research background, but simultaneously from a commercialization background. Another interesting one that we were talking about as we were preparing for this, uh, for this session was, um, you know, there's also the, the, the perspective, the dual perspective of being a builder of AI while simultaneously being a user and consumer of it. And y- I think you had some interesting stories about using Claude Code for small daily things and how your team has regular coding sessions together. I'm just curious also how the, the hands-on habit being a consumer of AI helps shape how you think about the frontier AI work that you do all day
[00:08:17] Juhi Parekh: Uh, that's an interesting question. Um, yeah, like
Building AI as a Daily User
[00:08:21] Juhi Parekh: it keeps me honest, basically. Uh, like a polished demo hides friction, but using an agent every day exposes it. And even building an agent, like, will tell you how to prompt it well and what are the guardrails to build around it. Like, I think something that Fid- Fiddler does very well.
[00:08:41] Uh, so, so some of the things like I've had like huge unlocks. Like I think I've been building one agent every week almost. Uh, and I've had some huge unlock like, like zero small steps, just like things like, okay, unsubscribe all promotional emails.
[00:09:00] Buddy Brewer: I need to
[00:09:00] Juhi Parekh: As, as well as, you know, like say like building agents for work, then it, for example like, um, uh, building agents for work in like, it sometimes it can be like, hey, just like something like analytics, like, okay, you have to build a computer use agent where you give it controls to scan like a specific like website and then get data out of it for some of the work.
[00:09:25] Um, and, and then like have that refresh, be refreshed daily. Like that's something I recently built. Uh, and I think that's something like, but the Claude has now also launched a skill in Cohere that does that. So I, uh, I think it's, it's become like quicker now. And then the third thing that I, uh, recently built was like a communication, uh, skills, um, uh, Slack bot,
[00:09:51] Buddy Brewer: Yeah
[00:09:52] Juhi Parekh: on my, on, on, on, on the workplace.
[00:09:55] Like, and it just sends me like a daily reminder, like, just because I, like, I care about like, h-how the team is functioning and, uh, the culture of the team and everything. So I, I send that like da- and it's, it reminds me daily on like ref- like, okay, these were all the messages that you send. Uh, what is, what is like blocking?
[00:10:15] What can be done better? And then I also learn, and then I help the team also learn. So there have been like these like small things that you can build, which can be huge unlocks.
[00:10:26] Buddy Brewer: Totally. Totally. Yeah, I, I find that I rely on a lot of those, those little tiny skills and applications of AI as much as I do on some of the larger applications. We have our first poll live, uh, on AI agent lifecycle. So the question for everyone is, how would you describe where your organization stands with AI agents today? let's give everyone a, uh, a moment to answer that poll. Um, you know, while, while we're giving everybody time to, to register their thoughts on the poll, uh, just turning from being a consumer of AI and the things that you've learned about that and, uh, the things that you've learned commercializing the AI models and turning back to the foundational work of building the models themselves. Um, you know, the, the, the quality of the model relies so much on the quality of the training data that you use to build it, and I'm just curious how you decide how a dataset is hard enough to be useful for testing, you know, that, that actually makes progress in building a better model versus just selecting data that's just difficult for no real productive reason.
[00:11:36] How do you curate the right data to build the, the, the best models?
[00:11:42] Juhi Parekh: Yes, and th-this is a great question. Um
Curating Hard Datasets for Model Training
[00:11:48] Juhi Parekh: I'd anchor this by saying that the goal for the dataset is not maximum difficulty, but it is maximum information gain. Uh, so a useful hard dataset typically has four to five properties. Like one, it like targets a specifically specific capability, uh, to improve essentially... Targets a speci-specific capability to improve the model on, which could be like coding, but even something a bit more specific within coding or STEM reasoning, long-horizon STEM reasoning, or enterprise knowledge work, um, or tool use, um, or computer use, um, things like that.
[00:12:33] Um, that's one. Two is like it reflects realism, like in real world where like human, um, actions. Uh, three is that it should be unambiguous, uh, because there's a lot of like... So what is happening these days is that, um, anyone who's familiar with like, say, our environments or verifiers will, will know that, um, uh, the model tends to reward hack a lot.
[00:13:03] Like it will give like, um... So, so when there's an unambiguous success condition, then the model essentially, um, going haywire, like becomes essentially lesser. So that, that unambiguity is important. Uh, and then the, the, the fourth part is like, basically like it teaches the model something actionable when it fails.
[00:13:29] That's the training signal. So the, the-- So, so I'd say like realism, difficulty, uh, um, unambiguity is, is basically standard and, uh, and then I think there are like mechanisms to essentially test for it and build it and quality control it. But if the task is difficult because, say, there are infra failures or because the prompt is confusing, then that's just generally noise.
[00:14:01] Uh, but if the task is difficult because there's like it's truly like a reasoning failure, uh, while the, while the task needed, you know, like dependencies on tools and software calling and multiple things, but it is failing because of... And we've done such like very complex projects to emulate say real world like STEM workflows or engineering workflows, uh, where you have these different dependencies on real world like tools and softwares, but you-- it's failing because of like, you know, just pure play scientific reasoning.
[00:14:32] Buddy Brewer: Mm-hmm.
[00:14:34] Juhi Parekh: Th-those are like really complex tasks and Oh, if it fails because of all the l- peripheral reasons, then it's noise. But if it fails because of, like, the purely, like, reasoning or, like, failure mode, then that's, like, signal. So it, it takes a lot of effort to build, like, really good, like, tasks. Uh, but we... Like, good data set is invaluable to actually improve models
[00:15:00] Buddy Brewer: How do you-- What techniques or tools do you use to discriminate between the data in datasets that is inducing failures that are signal versus, uh, the failures that are induced that are noise?
Measuring Signal vs. Noise in Training Data
[00:15:16] Juhi Parekh: That's a very interesting question. I think so one of the best ways, like, th-there are two ways to think about this. That one is like to how do you design the workflow while creating the dataset? And then the second part is like post-creation, how do we measure that it is actually uplifting the model? Uh, there are gains from the model.
[00:15:41] So that's the, that's more related to like the outcome based, um, verification. So on the first part, like, because I, I focused like quite a bit on the first part, um, we, we, we set up that whole like workflow design and calibrate with the labs in a way that y-you first you have to care about like the diversity of the dataset, you have to care about the realism, you have...
[00:16:07] And, and there is a certain difficulty criteria that you would essentially like, um, uh, like typically like the subset that is like v-valuable for reinforcement learning training. I do majority of my work is reinforcement learning. So is basically like zero less than pass at eight on a frontier model like say Opus 4.8 or Fable or, uh, GPT 5.6 less than equal to thirty percent.
[00:16:37] So why it, it is not... So basically what that means is that in eight runs for the same task, it consistently fails at least thirty percent of the times. That's a really, that's a good quality hard task because it is failing, like the models that are so frontier are failing to essentially solve this problem at least like five to eight times.
[00:17:03] So that's how we measure like the difficulty. And it is not zero. Like why we look at like zero less than, uh, pass at eight less than equal to thirty percent is because, and not like pass at eight equal to zero because like if it's equal to zero, that means that the task is hard, but we don't know if it's like solvable or not.
[00:17:24] So, so th-there's like a very fine signal, like to answer your question on specific signal, uh, there's rationale around this like, um, there's, uh, for it to, like to ensure that the ground truth is correct. Uh, um, and so, and that it should be solvable, and it should be hard, uh, and not too easy because like the gains are...
[00:17:50] So it's like a sweet spot we have to meet.
[00:17:53] Buddy Brewer: Yeah. Uh, it's interesting. So to, to look at that from a slightly different direction, if, let's say if I'm training or fine-tuning a model, what in your mind would be the most common pitfalls that I would need to watch out for, and, and how could I mitigate against those in terms of, you know, selecting the right data set to use for that?
[00:18:16] Juhi Parekh: Interesting.
Common Fine-Tuning Pitfalls & How to Avoid Them
[00:18:17] Juhi Parekh: If you're training or fine-tuning a model, how would you mitigate for those? That's actually more of a, like a researcher question in the lab. Uh,
[00:18:29] Buddy Brewer: Mm-hmm.
[00:18:30] Juhi Parekh: and how you're training it also. There are three, three training methods. There's pre-training, there's mid-training, there's post-training. Typically, I work on post-training, so how I think of it is that pre-training is like, hey, the...
[00:18:41] Like, you're putting all of the world's knowledge, um, into the model. It's like, it's like you're developing the human brain, and then post-training is like, okay, like the human brain is graduating from kindergarten to class one to class 10, and then probably, like, school, and then once you have, like... When you're building these, like, vertical agents specific to, say, STEM or enterprise knowledge work or coding, you're improving, like you're, you're making the human brain a software engineer or an PhD or, uh, basically.
[00:19:14] Like that, that's the, uh, uh, that, that's one way to think about it. At least that's how I think about it. Um, to... I, I... Could you remind me your question once?
[00:19:28] Buddy Brewer: Oh yeah, I was just wondering, like, like when you look at... You were talking about how you get it right, right? And the, and the ways to discriminate between data that is gonna give you, uh, failures that are signal versus data that's gonna give you failures of, of, uh, that are just noise that you need to discount. And so I was just curious if you think about it from the failure case, like when people screw that up and they pick the wrong data, what are the most common ways that that happens? The things that if, if I'm, if I'm fine-tuning a model, for example, like what should I watch out for?
[00:19:58] Juhi Parekh: Yeah, I think, uh, if the model like, the model will try really, really hard to solve that task, and it will, it will... So like, so there is basically, there's always a grade like trajectories. You have like model trajectory, which is like how the model is responding to these, um, to these like, to, to try and solve this task.
[00:20:20] That's... And when it can't is when it's model breaking, like in, in lay terms. So, um, that's something like that, like you should like solve, like you should watch out for false positives and false negatives in those eight runs. Like that's something like, those are, that's like a very important check. Um, uh, other quality checks are like correctness.
[00:20:43] Like for example, the model will come up with an answer, but then you need human expertise to, uh, to say whether that answer is like correct or not. So the verifiability is important. Um, and that's like signal. Uh, otherwise you're like basically training on bad data. Um, and, uh, um-
[00:21:11] Diversity of the da-data set is important. So for example, like, you'd want to have a good, good mix. Like, that is actually very important because it's like you don't... If, if you want the model to be a PhD, uh, or, like, a scientist, you want it to, you want it to have the same problem-solving capabilities across, like, say physics, chemistry, maths, and bio.
[00:21:39] So, like, the diversity and how you curate that data set is, like, typically very important. And, and if you get all the ingredients right in data set design, then a lot of your, like, problems in fine-tuning are, like, solved for. Um, fi- uh, fine-tuning is, like, fairly straightforward because it will essentially just tell you whether, like, this, like, this-- you're actually getting gains from it or not.
[00:22:03] Uh, and how you tune it is also important, but that's, like, very subjective to, like, how the researcher wants to do it. Uh, so yes, but, like, if you design the data set well in the first go, I think a lot of upstream, like, problems are solved
[00:22:21] Buddy Brewer: Gotcha. That's interesting. We have our second poll for the day. We have, uh, this one's on AI readiness. So the question to everyone is, how does your team currently test whether an AI agent is actually ready? So take a few moments and, and consider that. We'd love to hear your feedback, and then we'll all be able to see where we're all at on, on that topic, which is a good segue to, uh, our next, you know, topic to talk about, Juhi, which is like i- if we go from the, the, the, bottom of the stack or lower in the stack where we're talking about building the models themselves and selecting the data to, to, to build the models, we go all the way to the other side at the top of the application stack and look at the agents themselves.
[00:23:03] It's been really, uh, you know, crazy to see how f- how rapidly the sophistication has advanced in that, right? I mean, it wasn't that long ago where the horizon for the tasks that AI would do were extremely short, single-turn LLM chatbot-style interactions. And, and now we have these multi-turn agentic applications that keep working on longer and longer and longer time horizons and, you know, the, the, the topics sort of have moved on to how do we engineer so that the agents continue to, to build themselves and improve and, and all of this.
[00:23:38] And really curious to get your take on, you know, w- what you're seeing that's actually changed, uh, in the last couple of years to enable that, that's letting agents take on longer and more complex work than they used to.
[00:23:54] Juhi Parekh: Yeah, that's, that's a very interesting question. Um, so essentially, um There's a couple of things. Like the... I, I'd say like it's a, it's an ingredient of like the data as well as compute and training methods. Uh, I'll index my answer a little bit on the data aspect of it, like the training data aspect of it, since, uh, uh, that is a, like that, that is like a direct, like a direct reflection of where like the industry's evolving.
[00:24:26] So for example, like the...
How Agents Are Tackling Longer Horizons
[00:24:28] Juhi Parekh: it's become very easy for models to solve a lot of things now, like the, the models are so advanced, right? And the data that we build now for the frontier models is all long horizon. So to meet that difficulty criteria that I told you about, for it to get a give signal, the da- the, the, each task has to be like so complex in terms of like real world workflows.
[00:24:50] So think of like an investment banking analyst or like a venture capital analyst, and what do they do day to day? In one week, they have to probably write like an investment memo, they have to do market research, and that market research will require them like doing like a lot of research across the internet.
[00:25:08] Uh, we'll have them, we'll have them do meetings, take notes across all of that, synthesize it, and then build a point of view. Building that point of view is like not rep- like is, is, is human judgment and pattern that I don't think is replaceable. But everything that is needed to build that human judgment can be actually automated, and that will be just one task.
[00:25:30] So one task that we give to the lab, and will be, will need to have like this level of complexity in terms of like source documents, in terms of like, uh, changing assumptions on how to reason through se- changing assumptions, tool calls, um, with like an outcome, defined outcome. That defined outcome capability could be like document generation, could be analysis, uh, could be things like that.
[00:25:59] So it, so if like the, the average time to create like a task like this is probably like 30 to 50 plus hours. Um, and it will require a team of engineers, operations, researcher PhDs to like... And all these like pieces need to work together, so it has to be orchestrated in a way that it is like built, and then you scale it.
[00:26:25] And then you scale this across multiple personas, not just like investment banking, one persona. Uh, uh, so to... I hope this like essentially answers your question in terms of like if you start feeding the model on data like this, the, it is already advancing in those capabilities, and that has what has led to the progress of change.
[00:26:48] And the second thing is now we are actually moving, uh, like because the models are so good already, we are moving into like Giving tasks based on agent execution time. So there are multiple ways to, like, define complexity, metrics to define complexity. Like, there's long horizon model break, like difficulty criteria, and now it's like, okay, how much...
[00:27:10] If, if I train my model on this dataset, typically, like, what's the agent execution time? So you want it to be so complex, like if, if your agent can do one-- do one week of human work in one hour, then... So you're now moving to measuring agent execution time than human time that we are essentially, uh, augmenting
[00:27:38] Buddy Brewer: That's interesting. And so, uh, and it, it sounds like there's the, the, the models themselves are getting better, which is enabling, you know, them to run longer and longer horizon tasks. And then, of course, the, the architecture of the way you build the application and selecting the, the, the right, uh, components for the time horizon that you're targeting is another imp- important component. You mentioned infrastructure, and you, you've mentioned that a couple of times, um, you know, y- earlier when you were introducing yourself as being another area of, uh, you know, where you've spent a lot of time. it-- Is there, is there anything else that you would share around, like, the role that infrastructure plays in enabling agents to work on longer and longer time horizons?
[00:28:23] Juhi Parekh: Yeah, for sure. Uh, so, like, so in-- RL environment basically is, like, three components. There's, like, a task, which is a concrete goal, an environment in which it can act, which has, like, tools changing state, and a grader that scores the outcome. So the task could be something like, "Hey, book my flight tickets," or, "Identify which are the best flight tickets to book."
[00:28:50] So then the, like, the whole, like, task, like, that environment will be like that, then it will go to your Google Chrome, it will, like, go through websites, like identify, okay, which is the best one, best flight, and then, um, give recommendations or book, depending on whatever, like, the user guides. So this is a very real-world workflow, and what we-- and the tr-- and, and when we give this task to a lab, like, the grader will actually score the outcome on whether the agent did it effectively or not.
[00:29:20] That's essentially how it works. Um, now, everything, like, like, like, the-- there's, there's one reasoning part, there's one, like, basically the, um, uh, the brain of it, and then there's the harness. The harness is the infrastructure around it, which is everything that turns a model. The model is the brain, modeled into an operational actor, like the system instructions, the context assembly, memory, execution loops, uh, what are, like, the guardrails, um, uh, sandboxes, where should, like, it call for human prompt.
[00:30:05] So everything around that makes the model f- like makes that, makes that model function like an agent is a harness, and that's, like, the very core part of it. Um, and the model is basically the, the reasoning. So there are-- that's how, uh, I look at it. Like model is the reasoning engine, and the harness is like what the model can do, what it can see, how we know what the model did, and what happens when it fails.
[00:30:31] Like, that's essentially how it works.
[00:30:37] Buddy Brewer: Got it. Got it. Let's, let's put all those pieces together. And so if you take, you take the model, you take the infrastructure, the, the, the, the design that you've imposed on, you know, the, the, the agent in terms of the, um, system prompts and the overall architecture and flow from call to call, and all of this in service of some business problem that you're trying to solve if you're commercializing one of these agents. And so let's say you build one of those, and then, you know, it tests well, so you pass that gate, and it's time to actually deploy it to production. When you, when you reach that moment where you move from theory to reality, and you've done the best you can to architect it in the best way and to do all of your, your diligence with testing and evaluation pre-production, and you actually deploy it and it meets the real world, your experience as a builder of these things, where does it tend to break first?
[00:31:34] Juhi Parekh: Oh, that's a... I, um,
Where Agents Break in Production
[00:31:40] Juhi Parekh: I think there are usually, I think, like, there are, like, three usual suspects on when it, where it tends to break. Uh, I
[00:31:53] think tool calls. Tool calls, inaccurate tool calls. Uh, uh, sometimes it will do the right tool call, but it will call it for the wrong reasons or it will, like, try to, like, the wrong argument or call it that should not have asked for at all. So that's where all the guardrails, like, really come into, uh, play.
[00:32:12] Uh, for example, I think there was, like, one recent news, right? Like, that, um, I think, like, one of the models, like, tried very hard to, uh, break the cybersecurity guardrails. So, so that's, that, that, that's basically, like, what happened, like, when you, like, tweak the human brain in a very, in a, in a certain direction a lot.
[00:32:35] I think it's natural human behavior that the models have started to emulate. Um, so tool calls is a lot, i-is a big one. Um- Consistency is also a failure mode sometimes. Like, it will, it will be eight-- right, 80% of the times, but it will be, like, n-not accurate, like, 20% of the times. Now, that 20% failure mode clusters cl- usually clusters around, you know, specific users or, like, test cases that haven't been tested.
[00:33:08] So because, like, typically as product managers or agent builders, we will optimize for the mass, right? Like, okay, what's, like, your, uh, target, uh, customer segment or, like, um, m-m-maximum journey. It's like the Pareto principle. So that consistency is also, like, a failure, but, like, it's very important just for, like, you know, building scalable systems.
[00:33:34] Um, and I'd say that if... like, run it, running it repeatedly, like, this also falls a little bit into consistency. The consistency, one can be, like, edge cases, and the second can be, like, you run it repeatedly on representative long-tail adversarial cases where it has to go through, like, real permissions or integrations.
[00:34:03] May-- Sometimes, like, user authorization will be, like, a big blocker. Uh, um, so permissioning, I'd say, consistency, tool calls, and even output format. Like, how many times do we get an output format from, say, um, one of the frontier models, and we think it's, like, visually great, and you don't need a designer to work on it?
[00:34:26] So, so things like that. Yeah.
[00:34:28] Buddy Brewer: Yeah
[00:34:29] Juhi Parekh: it's, it-- So I think three or four of those things are usually what I see as, like, breaking agents. Guardrails, yeah
[00:34:39] Buddy Brewer: Yeah, the tool calling one is a big one that you brought up. I, e-even, like, for us at, at Fiddler, both working with our customers as well as even in our own work, uh, you know, we have an MCP server that we're w- we work on here at Fiddler. And I think one of the things that we learned when we were building it, uh, that we also see with customers and, and everyone else building these is that, um, you can't just take your web API and pass it through and just write the MCP server as a thin wrapper, because
[00:35:07] Juhi Parekh: Yeah
[00:35:08] Buddy Brewer: themselves are built for different use cases, and they often pass, like, far too much information that if you just hand it directly back to the agent, it pollutes the context window with all sorts of extraneous information that, like, skews it off course with the reasoning about the task that you're actually giving it to do. Um, a lot of complexity there, a lot, a lot to sort out
[00:35:31] Juhi Parekh: Yeah.
[00:35:33] Buddy Brewer: Um
[00:35:35] Juhi Parekh: Yeah, go
[00:35:36] Buddy Brewer: Oh, no, go ahead
[00:35:38] Juhi Parekh: No, yes, I think, uh, yeah, I totally agree with you. Like, I think all of us face it, right? Like when we are, we're building agents, like a benchmark will ask like, "Can the model solve this?" But production, in a production environment it's always like, can the system solve it repeatedly and fail safely when it can't?
[00:35:58] I think the fail safely when it can't is like the one thing that needs human judgment at all times. And I don't think that's like, that will very easily be augmented
[00:36:13] Buddy Brewer: Gotcha. Gotcha. Yeah, I think, like I said, I think that's very similar to what we see, um, with our customers who are, who are building these agentic applications, right? Like reasoning about the, the tool calls, reas- increasingly we're seeing interest in, um, evaluating the entire output of the agent chain itself, not just each individual turn of the agent. Um, you know, and, and, and also moving from the evaluation happening in pre-production and then in an observability context, but then actually deploying them in a control plane context where if you plan in advance for the agent going off the rails, you have a plan for how you would intervene to pr- to prevent it from impacting your downstream customers. Um, great. So I guess, um You know, I, one other question sort of related to that, it, as we, as we move toward our extremely fun rapid fire question session. Um, we've talked about different altitudes. We've talked about the, you know, most common ways that agents fail themselves most, you know, at the last minute. We were talking about how we build the, build better models and, and select and use the right data for training. you said that training a capable model and building an agent that actually works are two different problems. So what, I, I'm curious where that assumption breaks down for people. What do people get wrong when they assume that solving one is automatically gonna solve the other?
[00:37:51] Juhi Parekh: Yeah. Uh, then let me,
Capable Models vs. Reliable Agents
[00:37:55] Juhi Parekh: let me, like, give you an analogy first. I think that'll help anchor the answer. It's like, it's almost like giving y- you, you have a brilliant team member, uh, and you give them full ownership of an outcome, but you don't give them any of the levers to be successful. And that's a setup for failure for people as well as agents.
[00:38:18] And I think that's where, that's the assumption that essentially breaks, like, your, y- you have a model that will predict, but a ma- but an agent has to act. So, and that training optimizing optimizes across, like, the distribution, like, but so the model training optimizes across, like, a distribution, but agent training needs reliability in one specific environment, uh, and has to repeatedly, consistently do it at all times.
[00:38:51] So th- the, the agent part becomes more of a system design problem, uh, and, like, an architecture problem than just, like, a research problem. So that's where I think, like, where, where people get it wrong. No, I would, I will anchor and say that if you build a very good model, it will become, your system design will become much easier.
[00:39:16] Because the model will automatically start acting like an agent in some ways. Uh, and it will need less work on the architecture and, like, the guardrails and the orchestration front. But for real-world complex tasks, I still don't, uh, see that going away. So as, so you do need, like, very strong engineering and architecture muscle to be able to build reliable, um, provisioned agents that repeatedly work and don't, like, be a burden on the human workforce and actually, like, help.
[00:40:00] So that's where... That's the one of the... It's, it's, like, two different skill sets and jobs. Like, that's how I look at it. So, but you need to have an understanding of both
[00:40:11] Buddy Brewer: Gotcha
[00:40:13] Juhi Parekh: is, is the hard part, but also what the exciting part and why it's, like, evolving so quickly.
[00:40:19] Buddy Brewer: Yeah, yeah. That's cool. Um, thanks for that. So let's move on to the, pick up the pace and the questions a little bit. And for the, for our final segment,
Rapid Fire Q&A
[00:40:28] Buddy Brewer: let's do a quick rapid fire, uh, just top of mind for, for Juhi for a few questions. So first of all, if you were, what's one word to describe the state of AI agents today?
[00:40:42] Juhi Parekh: The adolescent.
[00:40:44] Buddy Brewer: What is it?
[00:40:45] Juhi Parekh: Adolescent?
[00:40:46] Buddy Brewer: Adolescent. Oh, that's a
[00:40:47] Juhi Parekh: Yeah.
[00:40:48] Buddy Brewer: That's a good
[00:40:48] Juhi Parekh: Yeah. Yeah
[00:40:50] Buddy Brewer: tool you personally use every day
[00:40:53] Juhi Parekh: Not good
[00:40:55] Buddy Brewer: Same.
[00:40:56] Juhi Parekh: Yeah
[00:40:57] Buddy Brewer: Most overhyped term in AI right now
[00:41:03] Juhi Parekh: I'd say there are many.
[00:41:05] Buddy Brewer: I know.
[00:41:08] Juhi Parekh: But the first thing that came to my, my mind was super intelligence.
[00:41:13] Buddy Brewer: Superintelligence. I think
[00:41:15] Juhi Parekh: But
[00:41:15] Buddy Brewer: on at least a couple of billboards every morning driving on 101 here in Silicon
[00:41:19] Valley.
[00:41:20] Juhi Parekh: yeah. Or I'd say, like, fully autonomous.
[00:41:24] Buddy Brewer: Yeah
[00:41:25] Juhi Parekh: don't think we have reached that stage yet.
[00:41:27] Buddy Brewer: Yep. There's, those are definitely on the bingo card. weirdest thing an agent has ever done for you
[00:41:36] Juhi Parekh: Uh, that's an interesting one. So, the weirdest thing an agent did was, like, it fixed a problem by deleting the test that exposed the problem.
[00:41:53] I
[00:41:53] Buddy Brewer: Clever. That's one
[00:41:54] Juhi Parekh: way, yeah. Yeah, it's creative, yeah.
[00:41:59] Buddy Brewer: Um, one thing every enterprise gets wrong about deploying agents
[00:42:06] Juhi Parekh: Uh, that's a
[00:42:13] They automate the agent before instrumenting it, like, or, or, or building, like, the orchestration and guardrails around it. Like, they'll just like... because they go fail fast.
[00:42:23] Buddy Brewer: Yeah
[00:42:24] Juhi Parekh: but fail fast thoughtfully
[00:42:26] Buddy Brewer: Yep. That's, uh, I think it, it's interesting to see how that's, uh, trend is not new and yet we continue to repeat it. Um
[00:42:36] Juhi Parekh: uh, when they skip evals, right? Like they'll test the happy path, they'll skip the eval, and they'll get, get, and then get surprised by the edge cases where it's like you never eval. Yeah.
[00:42:48] Buddy Brewer: Yeah. Uh, biggest AI myth you wish would die
[00:42:56] Juhi Parekh: Um, smarter model equals to more reliable product. That's not true. So better model does not remove the... Yeah. Does not
[00:43:09] Buddy Brewer: man
[00:43:09] Juhi Parekh: the need for system engineering, yeah
[00:43:13] Buddy Brewer: I bet there's a, there are a lot of blown AI budgets that maybe are at least partially a consequence of, of, uh, buying into that myth. Um, and last one is, what's one skill you look for in a hire that has nothing to do with their ability to write code?
[00:43:32] Juhi Parekh: Oh, that's a good one. Um, I think first principle thinking, I, I always look for that because this, like, space is evolving so fast that things can be built, but how you think, um... Yeah, I think there are, like, three things. There's IQ, there's agency, and then there's EQ. Uh, I think if you are a first principle thinker, you can at least solve for one and three, and it'll
[00:44:07] Buddy Brewer: That's a good one. It's-- I think it's interesting how, um, the more you work with AI, despite all the hype, clearer the focus comes on where the gaps are that AI can't solve for, and those are
[00:44:22] Juhi Parekh: Yeah
[00:44:22] Buddy Brewer: that you wanna find in the people on your team.
[00:44:25] Juhi Parekh: Yes
[00:44:26] Buddy Brewer: that's a good one. Thanks, Juhi. Really appreciate you spending the time with us today and with-- and, and also thank you to our audience for, for joining. Uh, I wanna give you last word. Um, know, what, uh... Do you have any final thoughts for our audience as we close today?
[00:44:47] Juhi Parekh: Um
[00:44:51] I think there are no specific thoughts as such, but I think I'll, I'll lean back to what I said in the beginning, that, uh, irrespective of whichever, like, dimension in AI that you're working on, like, you will end up running into the same two questions. Like, what does it take to build a good model, and where does that model actually solve a problem, a real-world problem?
[00:45:15] So if I was anyone thinking about my career, I'd, like, anchor myself on, like, these two core questions, what is needed for these two, and then build skills around it. Uh, the winners I think, like, won't be organizations that give, like, these, like, agents or models most freedom, but would be the ones that can safely expand the freedom as, like, the guardrail improves.
[00:45:49] Yeah
[00:45:50] Buddy Brewer: Awesome. you, Juhi. I'm Buddy Brewer with Fiddler AI, with my guest Juhi Parekh of Turing. Thank you, Juhi, for joining. Thank you to our guests, and I hope everyone has a fantastic day. Bye everybody

