In this episode of The Dialogue Architects, host Lauren Goerz, Staff Product Marketing Manager at Rasa, sits down with Soo‑Jin Lee—a New York–based playwright and conversation designer—to explore how creative storytelling, metadata, and human evaluation shape modern generative AI systems. Soo‑Jin shares her path from theater to tech and breaks down how designers guide model behavior through prompts, guardrails, and structured evaluation. This conversation is essential for anyone building user‑centered AI experiences.
What can a playwright teach us about designing generative AI? In this episode of The Dialogue Architects, host Lauren Goerz welcomes conversation designer and playwright Soo‑Jin Lee for an in‑depth look at the creative and technical practices that shape AI behavior.
Soo‑Jin traces her journey from public‑school teaching and playwriting to building conversational search experiences powered by language models. She explains how modern conversation design goes far beyond scripting dialogue—instead involving system‑prompt design, content guardrails, metadata strategy, and rigorous human evaluation to ensure responses are accurate, safe, and aligned with brand values.
Using “smart search” as a real product example, Soo‑Jin breaks down how metadata underpins natural‑language queries and how short generated confirmation blurbs help users validate intent. She also demystifies the world of human raters, sharing what it takes to calibrate feedback, avoid bias, and make ethical decisions around inappropriate or sensitive content.
The conversation highlights the importance of collaboration between designers and engineers, and why hybrid creative‑technical thinkers are becoming essential in the GenAI era. Soo‑Jin closes by exploring the challenge of teaching AI humor—and her interest in designing bots that bring wholesome joy to user interactions.
This episode offers actionable insights for designers, engineers, and product leaders building trustworthy, user‑aligned generative AI systems.
Episode Timestamps
(00:00) Welcome and guest introduction
(00:47) Origin story: playwriting to AI
(03:29) Teaching models like students
(05:18) GenAI workflows and guardrails
(09:07) Smart search interfaces explained
(12:38) Personalization and metadata
(15:47) Human raters and calibration
(18:24) Ethics and brand‑safety decisions
(20:01) Feedback loops and RAG workflow
(21:25) Designers and engineers working together
(23:27) What humans create that AI can’t
(24:42) Future directions: humor in bots
(26:54) Wrap‑up and thanks
Soo‑Jin Lee is a New York–based playwright and conversation designer whose work blends creative storytelling with practical AI design. Before entering the world of conversational AI, she spent years as a public‑school teacher and playwright—experiences that sharpened her understanding of dialogue, audience intention, and the nuances of human communication. Today, Soo‑Jin works across generative search and language‑model‑driven experiences, where she develops system prompts, designs content guardrails, and leads human‑evaluation workflows to ensure accuracy, safety, and clarity. Her multidisciplinary background allows her to approach AI design with both imagination and rigor, making her a distinctive voice in the evolution of generative AI and conversational UX.
The Dialogue Architects
is a podcast from Rasa, hosted by Lauren Goerz, exploring the craft and strategy behind designing conversational AI in the enterprise. Each episode brings together technologists, designers, and product leaders to unpack how dialogue is built, scaled, and governed in real‑world systems.
Lauren Goerz: [00:00:00] Welcome to the Dialogue Architects, where we explore how enterprises can thoughtfully design, scale, and govern conversations between humans and machines. Today we're joined by Sujin Lee from New York, a playwright and conversation designer who has worked at the intersection of AI search and human judgment for one of the largest streaming providers and language model providers.
That we know and love today. She's worked on both classic conversation design and generative conversational search, and she's known for arguing that the real skill in AI isn't just prompting it's taste, judgment, and creative constraint. We're talking about the art and science of making search better.
Soo-Jin welcome to the podcast.
I knew you're doing lots of exciting projects right now, but I'd love to start with a little bit of an origin story. Tell us a little bit about where your career began and uh, how you ended up where you are growing up. I think I was always a word lover. [00:01:00] I enjoyed stories. I was a kid who always had the nose stuck in a book.
I just followed my love of narrative, language, character, all of that. And by the end of my college. Courses. I took a playwriting class and that changed my life. I found my format, I studied playwriting in grad school. 'cause I felt like, Ooh, I gotta do bootcamp. And if I wanna be a theater professional, I have to.
Prove it, you know, so you, you go get schooling, validation, all of that. And after that I kept plugging away with writing plays, but then picked up temp jobs, this and that. I landed in conversation design three years ago, right after the pandemic, but. During the pandemic and pre pandemic, I was teaching public school.
I heard about conversation design from a good friend. We were just having coffee and she told me how she pivoted in, discovered it and told me to do it. 'cause she's a fellow playwright. I'm a playwright. She's like, [00:02:00] you can do it. It's so easy. You know, like you have the training for it. Hop in, don't be afraid.
And I was like, I knew chatbots existed. Had no idea. That it took a lot of work. 'cause this was back in the, you know, N-L-U-N-L-P, the legacy way of doing chatbots and not the generative AI way. But that's when I joined and I think I'm grateful I joined pre the chat GPT explosion so that I knew the dialogue tree way, the old school way of how to build.
And then I'm still doing it, but. I have, uh, moved into the gen AI space because obviously that is a hot space and you want to stay relevant, although you can use both, both the legacy style, a lot of companies do hybrid anyway depending on what industry you're in, but that is me. The one through line I would say is language, love of language, love of character, so it's always using words to solve problems.[00:03:00]
I love that. I think that's one of the, anecdotally, one of the things I've noticed going to different con conferences in the conversational I or broader kind of conversational interface space is that there are a lot of theater kids, and I love it because, and I always figured it was maybe something to do with the idea of script writing, which you've done quite directly and how it overlaps with.
Conversation design kind of as a profession, but ultimately as you, as you mentioned, there's a lot more unscripted design happening right now, and I think that's what I'd love to learn a little bit more from you. You talked about teaching, working with students, however, you're still teaching, but not necessarily a student.
What does that look like in your career today? Okay, so I almost feel, I've heard this from another colleague, that working with models and teaching models to speak the way we want them to. That's like teaching a toddler how to speak, like use these words, not those, and giving maybe a reason why and it's shaping their branding.
[00:04:00] So. I think modeling, it's so, so funny that the humans have to model for the models so that they'll pick it up. So there are different tricks to doing that. But back to what you're saying, you know, it's guiding them to use language, the want, the way you want them to. It's the way what I was doing in the classroom as well with the students, like, Hey, this is an okay essay.
How do we make it better? Or, I always like to show the version that's not great, and sometimes it's the first draft, and then we look at like the polished one. I'm like, what's the difference? Why do we want the polished one? Why do we care? You know? And a lot of times kids might not be able to explain why B is written better than a, but.
They, I don't know. They all kind of know. And so you get that vibe, the feeling that it's, it's not as cluttered, it's clear language, or maybe even simple as elegant. [00:05:00] It is very much, I think a space where the humanities and science combine where we're currently in today. But I'd love to learn a little bit more about your day-to-day work.
'cause I think you have a really interesting job and I feel like there's probably gonna be more of you in the future. But what are you doing day to day? What are you working on? Then we can kind of ground and explore the space a little bit more. People are like, oh, conversation design is dead, or It's evolving, et cetera.
And even in the next year or six months, our titles are, our titles gonna change or, or what else are we gonna be responsible for? And then of course. We have engineers who will help build, and what's kind of happened, I think, in the gen AI space is that you really have to work closer together. But in a lot of companies, you're almost in separate rooms still.
You know, it's like be in the same kitchen and cook together, right? Even if your skills are. Your focus is, you know, a little, it's really together. It should be together versus separate. Um, I almost feel like it's having a playwright who writes the [00:06:00] script and the director who talks to the actors and other designers to put it on the stage.
You know? Mm-hmm. And if they work closely together, you see a play. We're launching a product together. And so the designers have a dream view and this is what they recommend. 'cause it'll be good user experience. And then you have the engineers who need to also have buy-in, get, you know, you get their buy-in and they're like, okay, we'll build it.
'cause if they're not interested, they'll be like, no. And often I feel like many of the companies be, be, you know, behind the scene. There are certain groups who have more power and get to decide where to steer the ship. But back to your question, um, conversation designers in the past, I feel like they showed, uh, show the, uh, workflow when this piece of script happens, this is how it's gonna answer.
And, and then, you know, you keep dividing into multiple, it's like the spider web of conversation and, um. What's interesting is it, it, [00:07:00] it wasn't that complex unless you take years and years to build that flow out. Mm-hmm. Or all the flows out. Uh, now it's almost like with the, if you use a model, unless you put the guardrails in, like the conversation can go anywhere and that's almost, um.
Exciting, but scary. And so, um, my experience now as a designer, I am, um, I help like mold, like system instructions. Uh, sometimes that's called the system prompt. Uh, and each company has their own vocabulary for all of this Gen AI stuff, which is like fun. But, uh, looking at that, and if we tweak this. What happens to the output?
So once our users engage with. Our interface. And that could be a chat bot, that can be voice assistant. How does our model reply? Is it being [00:08:00] accurate? Is it delivering what the user is asking for? Is it understanding what the user's asking for? And so, uh, for me, I'm helping with kicking off evaluations. So often I'm in the room with engineers who I often wonder, oh, who are they talking to, to, to get the guidance, who's steering their ship?
But. Often it's, I feel like it's a lot about the launch date. I gotta test, test, test to make sure we're not gonna be embarrassed when we push out our product. 'cause we, we want our customers to use the product. Right? So if it's not delivering good results, then uh, then yeah, then yeah. So, so for me, I feel like I am like a teacher where I am helping with instructions and often it's instructions for human evaluators.
To rate, what are we trying to see? How is the model performing? If we tweak this or that, whatever the, the recipe or things engineers, the builders are working with, [00:09:00] they'll tweak something or they'll push out a slightly different model. So a lot of times in gen AI space, it's a lot of testing, a lot of iterating.
So I guess a question just to kind of ground this in terms of the type of conversational interface that you're building, I think a lot of. People that come from either the conversation design space or the broader conversation AI space now called agentic ai, ultimately think of the work that's being done to help a brand with something like a customer service agent.
Ultimately, you know, something, a conversational interface that helps with customer service. You're actually working in a broader category, I would describe it. Um, you're working on actually the kind of. The models themselves, the ones that we, you know, the chat GPTs of the world, um, and also with some big brands.
And if I recall, it sounds like you're working on some of these publicly accessible ones and also some in-house ones. You have experience as well. Can you walk me through kind of the, the type of conversational interfaces that you're doing this work for? [00:10:00] Right. So, uh, I, I have, uh, dabbled in various interfaces.
Um, they go from something simple like searching for a movie. I want a certain movie and I wanna watch it, and you're on a streaming service, but it's helping folks find what they wanna watch, where the user uses natural language. So that's kind of cool. Instead of, you know, in the olden days, you have to make sure everything's spelled correctly, and sometimes the director's name can be very long, et cetera, or you forget.
You know, the actor's last name, maybe you remember the first name, et cetera. So, but what's cool is that with natural language, it's a little more forgiving and the, uh, model will help you find it based on what you have. It's almost guessing what you want, but you could describe it. Uh, so that is like smart search, if you will.
And, uh, and, and it's to eventually, the point is [00:11:00] to have something entertaining to watch and maybe, got it. You don't wanna be entertained and you wanna watch something that's documentary. That could be entertaining as well, but like, that's like more nonfiction. Uh, but that's something that I've, uh, worked on and I think people don't realize that that's, that can be AI as well.
Yeah. Most people think, well, you gotta talk to a robot. You know, and it's not just that we've had, you know, um, AI around for a while, we just didn't call it that or. Weren't aware, uh, the technology behind the scenes, uh, other interfaces, voice assistant, the smart voice assistants, uh, now in the past, so in the past we've had voice assistants where somebody behind the scenes had to think in advance, Ooh, I bet the user's gonna ask this question 'cause today's a holiday, or whatever.
And so they have, you have to, it's like you have to predetermine. What users will ask and then build that flow. Because if you don't build it, uh, the voice assistant will be like, I don't know [00:12:00] how to help you. And then the user's gonna be dissatisfied. So you need to have like a really, uh, strong data so the solution written somewhere so that the bot can go grab it.
So I'd love to explore a little bit more. You talked about different types of, so of course classical customer, um, customer facing, um, customer service agent. That's something that we at Ros of obviously we're, we're working pretty closely on. But I think also components of that often have the smart search element to it.
And I'd love to dive into that side. 'cause I think that's a really interesting space that, that has a lot of potential in it then that you're exploring right now. Especially as you talked about streaming services. I mean, at the end of the day, it sounds like then a customer comes to the streaming service and says something, Hey, I'm looking for this in that movie.
I really like documentaries in the crime genre. You know, something like that. And then, and then do, are you actually then generating that output, something that's personalized just for them? Yes. We want that to [00:13:00] happen. So in the backend, um, we, uh, there are like multiple, like. Uh, I would like to call them, I guess, fancy search engines, if you will.
And this is where you, you need the data to be tagged. Well, you know, so whatever that crime documentary you're talking about, it should have. Multiple meta tags like a lot. And so if that is done well, it's like a good library system. Then when you, um, send, you know, um, ask the search engine to find it, it will find, uh, the possible titles for you.
And then also we need language that's generated fresh for that query. And you get to, um, answer. So. The it, some people are like, do we really need the word blurb? Nobody reads anymore. But at the same time, the word blurb lets the user know, I understand what you're [00:14:00] asking for. Kind of like when you go to a restaurant and you order food and if the server repeats back your order, even if you're the only one dining, it's just, uh, verifying this is what I heard.
Are you, is this accurate? 'cause they might've, you know, I don't know, misheard, et cetera. Any, anyone can make mistakes including machines. So. Um, I think when and, and if the, uh, blurb is short saying, you asked for such and such and here's what I found. Right. Uh, that can be helpful Again, um, I wonder how many people read the blurb.
Some people love to read instructions, other people don't, and they just dive in and start building, you know, when you get the furniture set at home or whatever. Yeah. Love that. I'd love to dive in a little bit more on this. Personalization element. And I, I understand also in the course of this project, that means you're also part of the team that sounds like is labeling rating and improving the type of blurbs that come [00:15:00] back?
Do I do, did I understand that correctly? You know, it, um, it's about looking at the accuracy. Mm-hmm. Uh, most companies have lots of teams and so you have the team that is responsible for owning the meta tagging. Mm-hmm. And they're really good at that. Who is the audience for? Can children watch it? So that's also important, uh, tagging for that.
Who is your audience And, um, also, you know, maybe the emotions that get aroused when you watch certain things. Uh, and then of course you have the genres like thriller versus, you know, fantasy versus, so there are just so many, and I think the, the more language you use, uh, I don't know, there's a stronger chance of people getting what they want.
Kind of thing and you're on the team then that is rating or, or trying to improve the accuracy of the output. Now this is tricky 'cause I would imagine it's subjective and you need lots of different types of people with different backgrounds to kind of [00:16:00] meet the need of whoever your target customer is.
I'd love to learn a little bit more about this and what this looks like and, and your process and, and the types of teams that you work with. So I think what's important is to have, uh, consistent raters. And often I think, uh, even big tech companies, what they do is they outsource the human rating thing.
So it depends. Some companies are like, actually, let's have them in house. So we have them. And then having them consistently over, you know, multiple batches throughout a year, then we can really track it is our model. Um. Better at creating language that represents the brand. As well as being accurate and not offensive.
And you know, there are all these things you have to track, but I think if you use the same Raiders and they start calibrating. And so, um, that's an interesting thing. And I had to do rating myself as the [00:17:00] classroom teacher who taught it, and I had to calibrate with my colleague who was also teaching the course, you know, next door.
I don't know. When you outsource. The human evaluation if how to what extent that gets done. But if it's done in house, you can actually control that. And so, um, I've been in, you know, both scenarios and I feel like, uh, when you have the luxury of using the same raters. Um, that can be really, really helpful.
And then you gotta pick, um, sort of diverse Raiders, like maybe difference in background ages because, you know, um, entertainment platforms, uh, serve, serve all sorts of people and. You know, what might be offensive to one person, one reader might not be offensive to another. So sometimes you have to, you know, discuss uncomfortable topics in, in, you know, like how you and I are having a conversation right now, and it's okay to talk [00:18:00] about these offensive topics because people might search for this.
And there's that other question, do we return titles? You know, and obviously if we have the titles, yes. Otherwise we're blocking, we're gatekeeping, you know, titles we have in the catalog. But sometimes, you know, if people are asking for a certain title using inappropriate language, do we return it? Hmm.
Yeah. Huge ethical element there. Yeah, because branding everything. Yeah. Yeah. Because someone's gonna tweet about it the next day and then make the company look bad, and it's like, oops. So it's not just about saving the company's face, but also, you know, what does the company stand for? You know? 'cause it's always like, yes, it's fun to get people in trouble online or whatever.
Some people really are into that. Uh, but also as a company, where do you stand? And so it's a company's values and then. You know, you realize, wow, as designers, you really. Make a difference. Uh, and it's like AI ethics and practice. Yeah. When [00:19:00] you make those discussions and then, you know, most companies have legal teams who hover over this and guide you.
This is okay. Right now that is not okay. You know, because, um, a lot of, um, services now carry adult material and not just with cursing like, like rated r like that, but more adult stuff. And so. You know, um, how do you, yeah, it's like when people ask for certain things, how you keep online ultimately. Yeah. Oh, that, that's that too.
That is also tricky as well. But, um, yeah. Yeah. Okay. And also if you're looking at, uh, large language models, um, yeah. Do we really wanna satisfy the customer and exactly what they want and deliver what they want? Even if they're asking for things in a certain fashion? Yeah. So your company has to make a stand.
Like, yes, yes, we will be the best customer service ever. Or actually, um, we, we do have standards and these are our [00:20:00] standards right now, today. What does that actually look like? Are you combing through transcripts and saying, you know, thumbs up, thumbs down? Are you giving written feedback? Are you giving, what does that actual feedback loop look like for your team so that you can improve the results in the future?
Yeah. So it, you know, um, one way that has worked is that. You know, you plug into the model like you're mm-hmm. Sending like a customer query, like, Hey, I want to, you know, watch such and such a thing, and you get a response and we get to rate and judge how the model responds and do they use appropriate language, um, should this have been flagged so that.
We don't even answer. And they're fallback answers. Like, if you're asking for something really inappropriate, we'll say, you know, we don't carry that in our catalog. Try something else. Right? Something, something along that, those lines. Uh, but I feel like we're seeing how much is the model letting through?
Is it, [00:21:00] um, is it letting. You know, are we answering anything that's asking for adult material or are we blocking it and it's going to fall back? So we're judging that this has a parallel with rag systems in the broader kind of conversational ice space. Like what is, what is the workflow? Who are the personnel that are required?
Is that like, 'cause it sounds like you're working with engineers, but I don't know when they pop up into this kind of broader picture. Yeah. So, um, it varies. It depends on the company you're at. Some companies are, um, let you into the room with engineers. Uh, 'cause I think, you know, um, it's really like model UX work.
I feel like our conversation designers, in the past we were always under ux, but now since we're tinkering with models we affect, um. How the model gets manipulated. So, you know, we are in the [00:22:00] conversation and I think it's, um, companies who I think are wise will include the designers in a room with engineers.
And you might not understand 80% of what they're talking about. But then over time, through immersion, you like pick up the language. That's what I notice and that's valuable. And it doesn't mean you're back there tinkering too, but at least you know what's going on. You know, and it takes time to be like, Ooh, when they tweak this now, uh, the results are like this.
Let's not do that again. Or this model, this fresh model is actually delivering, it sounds like. So that, that means working with language models. One of the core things you've learned is like, basically you need to let the language people cook in the language sphere and you need to let the engineers cook in the engineering sphere, but make sure they're all in the same room.
Otherwise you're gonna, you're gonna end up with, with different and disparate user journeys. I think it would be a great collaboration and. Both teams can learn from each other, and it'd be great if there was like a hybrid person who could speak both languages. [00:23:00] They're gonna become really powerful, I think, and they might become the strategy people.
It's like speaking multiple languages. You can speak the design language, you can speak the engineering language. And the product language, and then you make your own company if you can do all of it. Right? Yeah. And I think those folks probably end up being like strategy people who can look from above and get the bird's eye view and know what the goal is.
But I feel that, um, the, there, there's that whole. Debate about what can AI not create that humans can? Hmm. And I feel like, you know. I think AI can be trained to be manipulative and get into like, ai, human romances or whatever, but doesn't really know what heartbreak feels like. Does it know what loss feels like?
Humans do. Humans get, you know, experience trauma, [00:24:00] um, they know what, you know, joy feels like as well. And so that's the thing. I think human experience is what creates art. And so I don't know if, uh. AI will quite get there. Although there are people out there who are like, yes, I believe there's consciousness, you know, in ai, et cetera.
I haven't met that model yet, but maybe, maybe that model exists already and just hasn't been pushed out to the public yet. I have no clue. Yeah, yeah. No, but I think that may, I mean, you see it in your own job that the. Uh, the simple distinction between right and wrong, which is entirely not so simple, is something that takes a team of people to, to determine and a brand guideline and a company stance and a viewpoint.
On that note, you know, kind of to close out today, I'd love to know. Where you'd like to expand your reach in the future. When you talk about, you know, taking on new and different types of projects, it sounds like you've already had an enormous wealth of diversity in smart search as you describe [00:25:00] it. Is this a space you wanna stay or are there other areas of the broader kind of conversational interface universe where you think, Ooh, I'd really like to go in that direction too.
So I'm thinking in two directions. I haven't given up on this, but I really love humor and, um, when I write plays. That's one thing I write, or poems I write very seriously in my mind, I'm serious. But when you throw it in front of a live audience, they're laughing. And some people might be like, that's great, that's great.
But like, that's not my goal. I'm not trying to be a comic writer. But it's naturally there. And I feel like, um, humor, even in humans, I feel like it's a rare thing. And funny people come from funny families and it's like this. Thing that gets passed down. And I don't think it could really be taught either.
And, and so if we could teach models to be inherently funny, I think it would [00:26:00] bring so much joy. Um, I have a boss, uh, who once said, you know, we can never have too many jokes 'cause like the world needs it. And these are, uh, not like mean jokes, but more of. Just wholesome, but not dumb jokes, you know? Yeah.
It's just, you know, you know, laughter, laughter and feeling joy from laughter. It's such a gift, and so I don't think people are focusing on. Funny bots as much, and that already sounds kind of cringe, but like a bot who really ha is a little sassy and that can bring joy to the customer service experience and, um, that's possible.
I feel like if someone can really tweak that or know how to bring that forward, that person will probably make a lot of money because it's not easy to produce. Thank you so much. I think it was such a good conversation on like. Smart search [00:27:00] human evaluators. I think that's a huge piece we haven't cracked yet in this space, what human evaluation looks like.
So it's a pleasure to learn from you. Thank you. Oh, thank you, Lauren. Thank you for this opportunity.