Scaling AI in Legal: Building Uber's Redlining Agent
Staff Software Engineer
Product Lead
Sr. Legal Director, Product Legal (Platform)
Introduction
Few domains demand as much precision and care as legal work, making it a hard test for putting an AI agent into daily practice. Beginning in early 2024 and continuing through 2025, before today’s agent harnesses were widely available, Uber built the first generation of our Legal Redlining Agent (LRA). It helps Uber Legal teams handle high-volume contract negotiations without compromising trust or compliance, and without replacing a lawyer’s judgment.
The most durable lesson wasn’t about the model or architecture; it was about bringing an agent successfully into lawyers’ day-to-day work. We started from the business problem: a steady stream of contract negotiations involving substantial repeatable redline work. We met lawyers where they already work, inside Microsoft Word®, so LRA could flag risky redlines and suggest edits within their existing workflow. We then iterated closely with lawyers on what the agent should do, piloting it with our Legal team in 2025 and folding their feedback back into the product.
In this blog, we share what we learned bringing that first generation into practice: framing the problem, earning lawyers’ trust, and improving the agent through their feedback. We close with how the first-generation system worked and how we would approach the problem today as agent technology has evolved.
On March 9, 2026, LRA was part of Uber Legal’s winning submission for Most Innovative Legal Department of the Year at the ALM Legalweek Leaders in Tech Law Awards.

Figure 1: Legal Redlining Agent in action.
The Problem: High Stakes, High Friction
Figure 1: Legal Redlining Agent in action.
At Uber, thousands of contracts are negotiated every year. Behind each one is a legal team reviewing and negotiating these contracts. This process can be time-consuming, repetitive, and often a bottleneck for deals, onboarding, and launches. It’s high-stakes, detail-oriented work where precision matters and delays ripple across sales, onboarding, and launches.
Yet, buried inside this high-stakes process was something surprisingly structured: similar clauses, repeatable redlines across clients, and standing legal reasoning patterns. In other words, there was a clear system.
Since the legal organization’s work is mission-critical to Uber, we saw an opportunity to automate repeatable parts and free up lawyers’ time for complex judgment calls, all without sacrificing quality or consistency.

Figure 2: Contract review lifecycle and average support times for enterprise deals.
The Solution: An AI-Powered Redlining Agent
Figure 2: Contract review lifecycle and average support times for enterprise deals.
The Legal Redlining Agent is an AI-powered smart assistant integrated directly into Microsoft Word as an add-in, where lawyers already work. It reviews client edits, understands intent, recommends policy-aligned responses, generates structured comments, flags risks, and continuously learns from feedback.
It combines:
- Document review to identify redlines (edits made by clients versus by Uber lawyers)
- Intent detection to interpret proposed changes
- Policy-aligned recommendations to accept, reject, or modify clauses
- Comment generation to explain Uber’s legal position
- Risk flagging to highlight clauses needing deeper review by the lawyers
- Self-learning feedback loop to improve future recommendations
Since launching with our Legal team, we’ve seen an over 20% reduction in average contract review time and 91% accuracy in AI-generated decisions. We’ve also received feedback like, “It saved me real time; especially with comment suggestions.”
What is the Legal Redlining Agent?
The add-in analyzes a Microsoft Word document and identifies any redlines from external parties. Using AI, it then processes these changes, attempting to understand the client’s intent and the legal implications. It compares the proposed changes against our team’s legal policies and guidelines.
Based on this analysis, the add-in generates a series of suggestions, which includes recommended actions, proposed comments back to the client that justify these recommendations with legal reasoning, and an assessment of the risk level associated with each change. Uber lawyers then review each suggestion, deciding whether to accept, reject, or modify it before final incorporation into the client response.

Figure 3: High-level workflow of the Legal Redlining Add-in / Agent.
Iteratively Improving the Legal Redlining Agent
Figure 3: High-level workflow of the Legal Redlining Add-in / Agent.
The system we just described didn’t emerge fully formed. Getting here required four major iterations, each teaching us hard lessons about what Legal AI actually requires.
Along the way, we used Uber’s GenAI Gateway, which allowed us to focus entirely on refining the product logic, working closely with subject matter experts.
Iteration 1: RAG
Our initial approach to the problem focused on developing a foundational RAG (Retrieval Augmented Generation) system. The Legal team provided existing playbooks and examples of negotiation turns. Our hypothesis was that by ingesting these playbooks to identify semantically relevant sections, and then feeding these sections along with proposed counterparty changes to an LLM (Large Language Model), the system would generate sound decisions.
Figure 4: RAG (Retrieval Augmented Generation) flow.
- Inaccurate semantic similarity. The nature of the “key” we aimed to embed differed significantly from the content we sought to retrieve, resulting in imprecise semantic similarity searches. This inconsistency led to unreliable application decisions.
- Inappropriate tone. During negotiations, the tone of a response is critical. The application frequently generated replies that were either overly defensive or excessively positive, leading to unnatural and unacceptable communication.
- Lack of generalization. The application produced effectively random responses when encountering negotiation scenarios not explicitly covered in the playbooks, limiting its utility in novel situations.
Iteration 2: Feedback
Recognizing the limitations of a playbook-only approach, we understood the need for the agent to be more adaptable and to learn continuously from real-world interactions. This led to the development of a self-learning loop powered by direct lawyer feedback.
Initially, our feedback mechanism was simple: we saved the lawyer’s decision (accept/reject/modify), the original text, the redline, and a thumbs-down signal. This provided a basic error signal but lacked important nuances.
To deepen the system’s understanding, we evolved the feedback loop to capture the lawyer’s expected action and specific comments. When we introduced the ability for the AI to draft modifications, we further expanded this to collect the final modified text used by the lawyer. This granular data proved to be a turning point. It enabled us to map the counterparty’s intent directly to the lawyer’s precise, improved language. The agent started getting much more accurate, adapting to the team’s preferred style and negotiation postures.
We also started surfacing relevant past feedback and rules directly within the suggestion interface. This gave lawyers transparency into why a suggestion was generated, building trust and allowing them to see exactly what influenced the AI’s output.
At runtime, the agent queries this feedback store, which now contains thousands of previous interactions. The system first conducts a similarity search with metadata filtering, narrowing the pool to about 20 pertinent documents. These candidates are then passed through a second filtering layer using state-of-the-art LLMs to verify alignment with the user’s and counterparty’s intent. The agent then analyzes the most relevant examples to decide whether to accept or reject the proposed change and generates a comment for the counterparty.
We also implemented an exponential decay weighting algorithm to construct the context for future prompts. This mechanism prioritizes recent decisions, ensuring the agent adapts to shifting legal stances and mitigates concept drift, while maintaining a balanced distribution of agree and disagree examples. This dynamic, balanced few-shot prompting allows the redlining tool to learn user preferences and policy nuances in near-real-time without requiring manual model fine-tuning.
Figure 5: Feedback–the first iteration.

Figure 6: Feedback–further refinement.
Iteration 3: Tone
Figure 6: Feedback–further refinement.
Addressing the persistent challenge of tone, we integrated an additional LLM call at the end of the processing pipeline specifically to modulate the output’s linguistic style. But we quickly discovered that getting tone right in legal negotiations is harder than it sounds.
Early outputs oscillated between two failure modes: sometimes overly defensive and adversarial; other times inappropriately accommodating and eager to please. Neither matched the measured, professional tone our lawyers use in actual negotiations.
The breakthrough came from a close collaboration with the in-house Legal team. Recognizing that tone is a critical component of legal communications, we empowered the lawyers to take direct ownership of the specific language defining the desired tone and conversational style. Our engineering team managed the anatomy, but the in-house Legal team managed the content of the three-part prompt anatomy.
Figure 7: A prompt anatomy that includes tone, style, and conversation starters.
Other techniques for reflection involve complex loops to achieve convergence on the desired behavior. We found that the presumption of bad quality doesn’t affect positive examples, but sufficiently corrects negative examples at a high reliability even with just one shot. We attribute this to the reduction in complexity, due to one less decision that the LLM needs to make in the single prompt. The lack of loops also limits token usage and latency.
Figure 8: Using reflection to improve tone.
Finally, we provide a section for Uber’s in-house lawyers to give conversation starter sentences as few-shot examples.
Within days, they’d tuned it to match their voice precisely. This led to a crucial insight: prompts directly influencing output format should be crafted and managed by the application’s subject matter experts or end-users with AI teams guiding to ensure that prompting best practices are met.

Figure 9: Feedback and tone improvements influenced by subject matter experts.
Iteration 4: Agentic Modifications
Figure 9: Feedback and tone improvements influenced by subject matter experts.
While determining whether to accept or reject a change is valuable, the real time-saver for a lawyer is drafting the counter-proposal. Simple generative text often hallucinates terms or drifts from the contract’s defined terms.
To solve this, we moved beyond simple text generation to an agentic workflow. When a MODIFY decision is reached, we reference historical feedback and rules to refine the clause.
This allows the tool to help lawyers draft high-quality proposed compromise language that reflects prior legal guidance and incorporates the company’s broader risk position for each particular issue.
Rules Database
Lastly, we realized that lawyers frequently work with their own templates, which are often shared across teams. This insight presented two advantages:
- Repeatable questions. There’s significant commonality in the questions lawyers receive, and repeatedly articulating Uber’s position with a consistent tone is time-consuming.
- Consistent search keys. Since the original document consistently uses the same language, we can generate reliable keys for searching when a document undergoes changes.
To leverage these benefits, we developed a rules database for lawyers. This allows them to select important or frequently modified text within an original template document and associate a rule with it. A rule encapsulates Uber’s stance, any fallback positions, and an example of the desired application response. When the application detects a change, it performs a semantic vector search on the modified sentences against our rules index. If a rule matches with high confidence, it’s injected into the context window. To ensure strict adherence, a final LLM step is incorporated into the workflow to verify that the final output aligns precisely with the retrieved specifications.
Figure 10: How rules processing fits into the overall picture.
Architecture
The architecture follows a thin-client pattern: the Word add-in handles document interaction and user input, while a Python back end orchestrates all AI operations. When a lawyer triggers an analysis, the back end runs a LangGraph™ workflow that parallelizes intent detection, risk assessment, and policy lookup. Vector stores (OpenSearch®) provide the memory layer, retrieving relevant rules and historical feedback to ground each decision.

Figure 11: Overall system architecture.
Microsoft Word Add-In: Meeting Lawyers Where They Work
Figure 11: Overall system architecture.
A critical design decision was building the assistant as a Microsoft Word Add-in rather than a standalone web application. Lawyers live in Word, so asking them to copy-paste between tools would kill adoption.
The add-in is built on React and TypeScript, running as a task pane alongside the document. It communicates with Word through Microsoft Office’s® JavaScript API to read tracked changes, apply modifications, and highlight clauses under review.
This integration wasn’t without challenges. The Office JavaScript API has significant limitations that required creative workarounds:
- Performance: The Office JavaScript API is inherently slow, requiring us to optimize at the application layer with aggressive caching and minimal round-trips to Word.
- API bugs: The tracked changes API has edge cases that cause crashes on certain document structures, requiring defensive loading strategies.
- Missing features: Deleted text in tracked changes returns empty strings, requiring us to maintain our own text state for accurate diff display.
These constraints shaped our architecture: a thin client that orchestrates Word operations carefully, paired with a back end that handles all the heavy AI lifting.
High-Performance Orchestration
We optimized for both latency and accuracy by parallelizing independent graph nodes. For instance, while the Intention Node analyzes the counterparty’s meaningful intent, the Risk Level Node concurrently assesses the clause’s danger profile.
Hybrid Decision Engine: Rules versus Learning
To balance flexibility with compliance, we architected a hybrid system that treats hard and soft logic differently.
For non-negotiable policies, we use a deterministic Rules Engine. Unlike standard RAG, this triggers only on strict semantic matches against a target sentence. If a rule matches, it acts as a hard guardrail, overriding the model.
For negotiable nuances (tone, strategy), we rely on the probabilistic feedback loop.
This dual approach allowed us to bootstrap the system, delivering high-confidence results from day 1 based on rules, while the feedback loop accumulated the data needed to handle more complex, unstructured scenarios.
Self Learning via Feedback
The core of our self-learning capability is a feedback loop that removes the need for manual model fine-tuning. We capture a comprehensive snapshot of every lawyer interaction. This structured dataset includes the original context (contract text, counterparty redline), the AI's analysis (intent, risk level), and most crucially, the lawyer’s full response: their final decision, any specific text modifications they applied, and their written reasoning. This granular data allows the system to model not just what was decided, but why, enabling it to replicate the nuance of senior counsel in future similar scenarios.
Time-Weighted Self-Learning
To prevent concept drift where the model might over-index on outdated legal positions, we implemented a custom feedback sampling algorithm based on exponential decay. When retrieving few-shot examples for the context window, we calculate a weight for each historical interaction.
Figure 12: Custom feedback sampling algorithm.
Data Storage and Vector Search
Our knowledge base relies on OpenSearch, utilizing nomic-embed-text-v15 embeddings for high-dimensional semantic retrieval. We maintain two primary indices:
- Rules index: Stores legal policies, vectorized by the target sentence.
- Feedback index: Stores historical negotiation data, vectorized by the counterparty’s intention rather than raw text, allowing us to catch semantically identical changes phrased differently.
When a new redline is detected, the system performs semantic vector searches against these indices to retrieve the most relevant rules and past decisions, ensuring the AI has the exact context needed to suggest the change.
Improving Rule and Feedback Retrieval Accuracy in OpenSearch
One advantage that we have in the problem space is that negotiation always starts from the same source contract. This means that we can anchor learnings and feedback to the same sections across multiple interactions over the same document.
First, we break down a legal contract into small chunks delimited by ';' or '.' or '\n' characters. Then, whenever the system encounters a redline, we perform a lookup on the original source sentence to retrieve guaranteed relevant historical occurrences, or explicit rules. Whenever a user interacts with the system and provides feedback, it’s the same verbatim source sentence that is used for ingestion into the database.
Figure 13: Leveraging OpenSearch to improve rule and feedback retrieval accuracy.
Session Memory
For the conversational interface, we use Redis® to maintain session state, allowing lawyers to ask follow-up questions about the document. Behind the scenes, we run continuous LLM-as-a-judge evaluations to monitor the quality of retrieved context (document_quality_metric) and the adherence of the model to company rules (rule_adherence_metric).
Today’s Approach: Harness Engineering
As part of our Agentic AI efforts across Uber, which includes Legal AI, we are exploring agentic harnesses such as Claude Code® and OpenCode® in conjunction with skills to structure and automate workflows. In parallel, we’ll explore incorporating legal ontologies and knowledge graphs to provide a semantic foundation for representing legal concepts and relationships, enabling more consistent interpretation and reliable reasoning.
Conclusion
Our journey through these iterations, from the initial RAG system to integrating user feedback, refining tone, and implementing a robust rules database, underscores that building a truly effective AI-powered legal assistant is an iterative process driven by continuous learning and tight collaboration with subject matter experts. Each challenge we encountered became an opportunity to deepen our understanding and refine the system, leading to a more accurate, adaptable, and user-centric tool that genuinely augments human expertise in complex legal negotiations.
Acknowledgments
We’d like to thank Alissa McDowell for her close collaboration and subject matter expertise, as well as Uber Legal, the AI Platform teams, and their leadership, Michelle Parker and Viv Keswani, for supporting this effort to bring AI to Uber Legal.
Cover Photo Attribution: Generated with OpenAI ChatGPT Images.
Claude and its logo are registered trademarks of Anthropic, PBC in the United States and other countries. No endorsement by Anthropic, PBC is implied by the use of these marks.
Codex and its logo are registered trademarks of OpenAI OpCo, LLC in the United States and other countries. No endorsement by OpenAI OpCo, LLC is implied by the use of these marks.
FastAPI and its logo are registered trademarks of Sebastián Ramírez in the United States and other countries. No endorsement by Sebastián Ramírez is implied by the use of these marks.
LangChain and LangGraph and their logos are registered trademarks of LangChain Inc. in the United States and other countries. No endorsement by LangChain Inc. is implied by the use of these marks.
Microsoft Word, Microsoft Office, Office JavaScript API, TypeScript, and their logos are registered trademarks of Microsoft Corporation in the United States and other countries. No endorsement by Microsoft Corporation is implied by the use of these marks.
OpenSearch and its logo are registered trademarks of LF Projects, LLC in the United States and other countries. No endorsement by LF Projects, LLC, the OpenSearch Project, or the OpenSearch Community is implied by the use of these marks.
Python and its logo are registered trademarks of the Python Software Foundation in the United States and other countries. No endorsement by the Python Software Foundation is implied by the use of these marks.
React and its logo are registered trademarks of Meta Platforms, Inc. in the United States and other countries. No endorsement by Meta Platforms, Inc. is implied by the use of these marks.
Redis and its logo are registered trademarks of Redis Ltd. in the United States and other countries. No endorsement by Redis Ltd. is implied by the use of these marks.
Austin Greco
Staff Software Engineer
Austin Greco is a Staff Software Engineer on Michelangelo, Uber’s AI platform, where he led the developer experience and helped shape it from the early days into the system powering machine learning across the company.
Meghana Somasundara
Product Lead
Meghana Somasundara is a Product Lead for the Uber AI Platform. Previously she worked on building AI Platforms with other companies.
Frank Tenente
Sr. Legal Director, Product Legal (Platform)
Frank Tenente is Senior Legal Director, Mobility & Platform Product Legal at Uber, leading legal teams supporting Mobility and cross-platform business units. A commercial and product lawyer, Frank has scaled B2B legal teams and champions GenAI to transform contract review and legal work.
Sean Po
Staff Software Engineer
Sean Po is Staff Software Engineer and the Technical Lead on the Uber Agent Platform team.
Rush Tehrani
Sr Manager, Engineering
Rush is an Engineering Manager on the AI Platform Team at Uber. He supports the teams responsible for deployment and serving of classical, deep learning, and generative AI models, machine learning on mobile, and generative AI API gateway.