Skip to main content
October 8, 2026

Scaling AI in Legal: Building Uber's Redlining Agent

Austin Greco

Staff Software Engineer

Meghana Somasundara

Product Lead

Frank Tenente

Sr. Legal Director, Product Legal (Platform)

2+
AI robot analyzes legal documents, highlighting text with a justice scale in the background.
Share this article

Introduction

Few domains demand as much precision and care as legal work, making it a hard test for putting an AI agent into daily practice. Beginning in early 2024 and continuing through 2025, before today’s agent harnesses were widely available, Uber built the first generation of our Legal Redlining Agent (LRA). It helps Uber Legal teams handle high-volume contract negotiations without compromising trust or compliance, and without replacing a lawyer’s judgment.

The most durable lesson wasn’t about the model or architecture; it was about bringing an agent successfully into lawyers’ day-to-day work. We started from the business problem: a steady stream of contract negotiations involving substantial repeatable redline work. We met lawyers where they already work, inside Microsoft Word®, so LRA could flag risky redlines and suggest edits within their existing workflow. We then iterated closely with lawyers on what the agent should do, piloting it with our Legal team in 2025 and folding their feedback back into the product.

In this blog, we share what we learned bringing that first generation into practice: framing the problem, earning lawyers’ trust, and improving the agent through their feedback. We close with how the first-generation system worked and how we would approach the problem today as agent technology has evolved.

On March 9, 2026, LRA was part of Uber Legal’s winning submission for Most Innovative Legal Department of the Year at the ALM Legalweek Leaders in Tech Law Awards.

Document editing interface showing tracked changes in a disclaimer text, with suggestions and options panel on the right.

Figure 1: Legal Redlining Agent in action.

The Problem: High Stakes, High Friction

At Uber, thousands of contracts are negotiated every year. Behind each one is a legal team reviewing and negotiating these contracts. This process can be time-consuming, repetitive, and often a bottleneck for deals, onboarding, and launches. It’s high-stakes, detail-oriented work where precision matters and delays ripple across sales, onboarding, and launches.

Yet, buried inside this high-stakes process was something surprisingly structured: similar clauses, repeatable redlines across clients, and standing legal reasoning patterns. In other words, there was a clear system.

Since the legal organization’s work is mission-critical to Uber, we saw an opportunity to automate repeatable parts and free up lawyers’ time for complex judgment calls, all without sacrificing quality or consistency.

Contract negotiation lifecycle showing steps: send template, client review, Uber responds, finalize contract.

Figure 2: Contract review lifecycle and average support times for enterprise deals.

The Solution: An AI-Powered Redlining Agent

The Legal Redlining Agent is an AI-powered smart assistant integrated directly into Microsoft Word as an add-in, where lawyers already work. It reviews client edits, understands intent, recommends policy-aligned responses, generates structured comments, flags risks, and continuously learns from feedback.

It combines:

  • Document review to identify redlines (edits made by clients versus by Uber lawyers)
  • Intent detection to interpret proposed changes
  • Policy-aligned recommendations to accept, reject, or modify clauses
  • Comment generation to explain Uber’s legal position
  • Risk flagging to highlight clauses needing deeper review by the lawyers
  • Self-learning feedback loop to improve future recommendations

Since launching with our Legal team, we’ve seen an over 20% reduction in average contract review time and 91% accuracy in AI-generated decisions. We’ve  also received feedback like, “It saved me real time; especially with comment suggestions.”

The add-in analyzes a Microsoft Word document and identifies any redlines from external parties. Using AI, it then processes these changes, attempting to understand the client’s intent and the legal implications. It compares the proposed changes against our team’s legal policies and guidelines.

Based on this analysis, the add-in generates a series of suggestions, which includes recommended actions, proposed comments back to the client that justify these recommendations with legal reasoning, and an assessment of the risk level associated with each change. Uber lawyers then review each suggestion, deciding whether to accept, reject, or modify it before final incorporation into the client response.

Word Add-In analyzes document edits using AI, rules, and feedback to provide recommendations, comments, and risk levels.

Figure 3: High-level workflow of the Legal Redlining Add-in / Agent.

Iteratively Improving the Legal Redlining Agent

The system we just described didn’t emerge fully formed. Getting here required four major iterations, each teaching us hard lessons about what Legal AI actually requires.

Along the way, we used Uber’s GenAI Gateway, which allowed us to focus entirely on refining the product logic, working closely with subject matter experts. 

Iteration 1: RAG

Our initial approach to the problem focused on developing a foundational RAG (Retrieval Augmented Generation) system. The Legal team provided existing playbooks and examples of negotiation turns. Our hypothesis was that by ingesting these playbooks to identify semantically relevant sections, and then feeding these sections along with proposed counterparty changes to an LLM (Large Language Model), the system would generate sound decisions.

Three-step process: Intention leads to Document Retrieval, which leads to Adjudication, shown with arrows.

Figure 4: RAG (Retrieval Augmented Generation) flow.

We quickly encountered challenges with this approach:

  • Inaccurate semantic similarity. The nature of the “key” we aimed to embed differed significantly from the content we sought to retrieve, resulting in imprecise semantic similarity searches. This inconsistency led to unreliable application decisions.
  • Inappropriate tone. During negotiations, the tone of a response is critical. The application frequently generated replies that were either overly defensive or excessively positive, leading to unnatural and unacceptable communication.
  • Lack of generalization. The application produced effectively random responses when encountering negotiation scenarios not explicitly covered in the playbooks, limiting its utility in novel situations.

Iteration 2: Feedback

Recognizing the limitations of a playbook-only approach, we understood the need for the agent to be more adaptable and to learn continuously from real-world interactions. This led to the development of a self-learning loop powered by direct lawyer feedback.

Initially, our feedback mechanism was simple: we saved the lawyer’s decision (accept/reject/modify), the original text, the redline, and a thumbs-down signal. This provided a basic error signal but lacked important nuances.

To deepen the system’s understanding, we evolved the feedback loop to capture the lawyer’s expected action and specific comments. When we introduced the ability for the AI to draft modifications, we further expanded this to collect the final modified text used by the lawyer. This granular data proved to be a turning point. It enabled us to map the counterparty’s intent directly to the lawyer’s precise, improved language. The agent started getting much more accurate, adapting to the team’s preferred style and negotiation postures.

We also started surfacing relevant past feedback and rules directly within the suggestion interface. This gave lawyers transparency into why a suggestion was generated, building trust and allowing them to see exactly what influenced the AI’s output.

At runtime, the agent queries this feedback store, which now contains thousands of previous interactions. The system first conducts a similarity search with metadata filtering, narrowing the pool to about 20 pertinent documents. These candidates are then passed through a second filtering layer using state-of-the-art LLMs to verify alignment with the user’s and counterparty’s intent. The agent then analyzes the most relevant examples to decide whether to accept or reject the proposed change and generates a comment for the counterparty.

We also implemented an exponential decay weighting algorithm to construct the context for future prompts. This mechanism prioritizes recent decisions, ensuring the agent adapts to shifting legal stances and mitigates concept drift, while maintaining a balanced distribution of agree and disagree examples. This dynamic, balanced few-shot prompting allows the redlining tool to learn user preferences and policy nuances in near-real-time without requiring manual model fine-tuning.

Flowchart showing user feedback storage, retrieval, and integration with legal redlining services and plugins.

Figure 5: Feedback–the first iteration.

Further refinement involved strategic data pruning to ensure higher accuracy in search results. This iterative enhancement enables continuous improvement in the application’s performance and yields a high-quality labeled dataset, crucial for tracking quality regressions and driving ongoing improvements.

Workflow diagram showing feedback processed through embedding, similarity search, LLM filtering, and result generation.

Figure 6: Feedback–further refinement.

Iteration 3: Tone

Addressing the persistent challenge of tone, we integrated an additional LLM call at the end of the processing pipeline specifically to modulate the output’s linguistic style. But we quickly discovered that getting tone right in legal negotiations is harder than it sounds.

Early outputs oscillated between two failure modes: sometimes overly defensive and adversarial; other times inappropriately accommodating and eager to please. Neither matched the measured, professional tone our lawyers use in actual negotiations.

The breakthrough came from a close collaboration with the in-house Legal team. Recognizing that tone is a critical component of legal communications, we empowered the lawyers to take direct ownership of the specific language defining the desired tone and conversational style. Our engineering team managed the anatomy, but the in-house Legal team managed the content of the three-part prompt anatomy.

Diagram labeled 'Prompt Anatomy' with sections: Objective, Style, Conversation Starters, each with related guidance.

Figure 7: A prompt anatomy that includes tone, style, and conversation starters.

The first part of the prompt we asked the lawyers to update is the objective. This section defines the role of the tone modulation agent. It implements a proactive and pessimistic AI reflection technique by assuming that the comment generated by the agent has the issues we had previously encountered.

Other techniques for reflection involve complex loops to achieve convergence on the desired behavior. We found that the presumption of bad quality doesn’t affect positive examples, but sufficiently corrects negative examples at a high reliability even with just one shot. We attribute this to the reduction in complexity, due to one less decision that the LLM needs to make in the single prompt. The lack of loops also limits ‌token usage and latency.

Flowchart showing tone transformation: negative and positive inputs become positive tone outputs.

Figure 8: Using reflection to improve tone.

Next, we prompt the node on the preferred style of lawyers. Since the agent is meant to provide the first pass of responses on behalf of the lawyer, we must ensure that the agent is direct, first-person, and doesn’t engage in idle conversation. 

Finally, we provide a section for Uber’s in-house lawyers to give conversation starter sentences as few-shot examples. 

Within days, they’d tuned it to match their voice precisely. This led to a crucial insight: prompts directly influencing output format should be crafted and managed by the application’s subject matter experts or end-users with AI teams guiding to ensure that prompting best practices are met.

User feedback flows through similarity search, LLM filtering, result generation, and tone adjustment using GPT 5.2.

Figure 9: Feedback and tone improvements influenced by subject matter experts.

Iteration 4: Agentic Modifications

While determining whether to accept or reject a change is valuable, the real time-saver for a lawyer is drafting the counter-proposal. Simple generative text often hallucinates terms or drifts from the contract’s defined terms.

To solve this, we moved beyond simple text generation to an agentic workflow. When a MODIFY decision is reached, we reference historical feedback and rules to refine the clause.

This allows the tool to help lawyers draft high-quality proposed compromise language that reflects prior legal guidance and incorporates the company’s broader risk position for each particular issue.

Rules Database

Lastly, we realized that lawyers frequently work with their own templates, which are often shared across teams. This insight presented two advantages:

  • Repeatable questions. There’s significant commonality in the questions lawyers receive, and repeatedly articulating Uber’s position with a consistent tone is time-consuming.
  • Consistent search keys. Since the original document consistently uses the same language, we can generate reliable keys for searching when a document undergoes changes.

To leverage these benefits, we developed a rules database for lawyers. This allows them to select important or frequently modified text within an original template document and associate a rule with it. A rule encapsulates Uber’s stance, any fallback positions, and an example of the desired application response. When the application detects a change, it performs a semantic vector search on the modified sentences against our rules index. If a rule matches with high confidence, it’s injected into the context window. To ensure strict adherence, a final LLM step is incorporated into the workflow to verify that the final output aligns precisely with the retrieved specifications.

Workflow diagram: Intention → Document Retrieval & Reranking → Adjudication → Tone Modification → Rules Processing

Figure 10: How rules processing fits into the overall picture.

When expanding this tool to additional lines of business, the rules database will provide a secondary benefit, helping lawyers maintain consistency in legal positions across the company.

Architecture

The architecture follows a thin-client pattern: the Word add-in handles document interaction and user input, while a Python back end orchestrates all AI operations. When a lawyer triggers an analysis, the back end runs a LangGraph™ workflow that parallelizes intent detection, risk assessment, and policy lookup. Vector stores (OpenSearch®) provide the memory layer, retrieving relevant rules and historical feedback to ground each decision.

System architecture diagram showing client workstation, backend service pipeline, and external services integration.

Figure 11: Overall system architecture.

Microsoft Word Add-In: Meeting Lawyers Where They Work

A critical design decision was building the assistant as a Microsoft Word Add-in rather than a standalone web application. Lawyers live in Word, so asking them to copy-paste between tools would kill adoption.

The add-in is built on React and TypeScript, running as a task pane alongside the document. It communicates with Word through Microsoft Office’s® JavaScript API to read tracked changes, apply modifications, and highlight clauses under review.

This integration wasn’t without challenges. The Office JavaScript  API has significant limitations that required creative workarounds:

  • Performance: The Office JavaScript  API is inherently slow, requiring us to optimize at the application layer with aggressive caching and minimal round-trips to Word.
  • API bugs: The tracked changes API has edge cases that cause crashes on certain document structures, requiring defensive loading strategies.
  • Missing features: Deleted text in tracked changes returns empty strings, requiring us to maintain our own text state for accurate diff display.

These constraints shaped our architecture: a thin client that orchestrates Word operations carefully, paired with a back end that handles all the heavy AI lifting.

High-Performance Orchestration

We optimized for both latency and accuracy by parallelizing independent graph nodes. For instance, while the Intention Node analyzes the counterparty’s meaningful intent, the Risk Level Node concurrently assesses the clause’s danger profile.

Hybrid Decision Engine: Rules versus Learning

To balance flexibility with compliance, we architected a hybrid system that treats hard and soft logic differently.

For non-negotiable policies, we use a deterministic Rules Engine. Unlike standard RAG, this triggers only on strict semantic matches against a target sentence. If a rule matches, it acts as a hard guardrail, overriding the model.

For negotiable nuances (tone, strategy), we rely on the probabilistic feedback loop.

This dual approach allowed us to bootstrap the system, delivering high-confidence results from day 1 based on rules, while the feedback loop accumulated the data needed to handle more complex, unstructured scenarios.

Self Learning via Feedback

The core of our self-learning capability is a feedback loop that removes the need for manual model fine-tuning. We capture a comprehensive snapshot of every lawyer interaction. This structured dataset includes the original context (contract text, counterparty redline), the AI's analysis (intent, risk level), and most crucially, the lawyer’s full response: their final decision, any specific text modifications they applied, and their written reasoning. This granular data allows the system to model not just what was decided, but why, enabling it to replicate the nuance of senior counsel in future similar scenarios.

Time-Weighted Self-Learning

To prevent concept drift where the model might over-index on outdated legal positions, we implemented a custom feedback sampling algorithm based on exponential decay. When retrieving few-shot examples for the context window, we calculate a weight for each historical interaction.

Exponential decay formula: weight equals exp(-lambda times age_days) in green monospace font.

Figure 12: Custom feedback sampling algorithm. 

Where λ corresponds to a configurable half-life (currently set to 365 days). This ensures the model adapts to shifting legal stances and concept drift without over-indexing on outdated precedents. Effectively, the system learns the team’s current preferences in near real-time, just as a human colleague would pick up on new norms.

Data Storage and Vector Search

Our knowledge base relies on OpenSearch, utilizing nomic-embed-text-v15 embeddings for high-dimensional semantic retrieval. We maintain two primary indices:

  • Rules index: Stores legal policies, vectorized by the target sentence.
  • Feedback index: Stores historical negotiation data, vectorized by the counterparty’s intention rather than raw text, allowing us to catch semantically identical changes phrased differently.

When a new redline is detected, the system performs semantic vector searches against these indices to retrieve the most relevant rules and past decisions, ensuring the AI has the exact context needed to suggest the change.

Improving Rule and Feedback Retrieval Accuracy in OpenSearch

One advantage that we have in the problem space is that negotiation always starts from the same source contract. This means that we can anchor learnings and feedback to the same sections across multiple interactions over the same document. 

First, we break down a legal contract into small chunks delimited by ';' or '.' or '\n' characters. Then, whenever the system encounters a redline, we perform a lookup on the original source sentence to retrieve guaranteed relevant historical occurrences, or explicit rules. Whenever a user interacts with the system and provides feedback, it’s the same verbatim source sentence that is used for ingestion into the database.

Three-step process showing a humorous contract, its simplified version, and a document labeled Rules & Feedback.

Figure 13: Leveraging OpenSearch to improve rule and feedback retrieval accuracy.

It’s important to achieve symmetry in query and input embedding to achieve the best search performance. In our case, the stored embedding and the embedded query are guaranteed to be similar as they are all sourced from the contract. Similarity search is still required because the document may change subtly with each round of lawyer annotations.

Session Memory

For the conversational interface, we use Redis®  to maintain session state, allowing lawyers to ask follow-up questions about the document. Behind the scenes, we run continuous LLM-as-a-judge evaluations to monitor the quality of retrieved context (document_quality_metric) and the adherence of the model to company rules (rule_adherence_metric).

Today’s Approach: Harness Engineering

As part of our Agentic AI efforts across Uber, which includes Legal AI, we are exploring agentic harnesses such as Claude Code® and OpenCode® in conjunction with skills to structure and automate workflows. In parallel, we’ll explore incorporating legal ontologies and knowledge graphs to provide a semantic foundation for representing legal concepts and relationships, enabling more consistent interpretation and reliable reasoning.

Conclusion

Our journey through these iterations, from the initial RAG system to integrating user feedback, refining tone, and implementing a robust rules database, underscores that building a truly effective AI-powered legal assistant is an iterative process driven by continuous learning and tight collaboration with subject matter experts. Each challenge we encountered became an opportunity to deepen our understanding and refine the system, leading to a more accurate, adaptable, and user-centric tool that genuinely augments human expertise in complex legal negotiations.

Acknowledgments

We’d like to thank Alissa McDowell for her close collaboration and subject matter expertise, as well as Uber Legal, the AI Platform teams, and their leadership, Michelle Parker and Viv Keswani, for supporting this effort to bring AI to Uber Legal.

Cover Photo Attribution: Generated with OpenAI ChatGPT Images.

Claude and its logo are registered trademarks of Anthropic, PBC in the United States and other countries. No endorsement by Anthropic, PBC is implied by the use of these marks.

Codex and its logo are registered trademarks of OpenAI OpCo, LLC in the United States and other countries. No endorsement by OpenAI OpCo, LLC is implied by the use of these marks.

FastAPI and its logo are registered trademarks of Sebastián Ramírez in the United States and other countries. No endorsement by Sebastián Ramírez is implied by the use of these marks.

LangChain and LangGraph and their logos are registered trademarks of LangChain Inc. in the United States and other countries. No endorsement by LangChain Inc. is implied by the use of these marks.

Microsoft Word, Microsoft Office, Office JavaScript API, TypeScript, and their logos are registered trademarks of Microsoft Corporation in the United States and other countries. No endorsement by Microsoft Corporation is implied by the use of these marks.

OpenSearch and its logo are registered trademarks of LF Projects, LLC in the United States and other countries. No endorsement by LF Projects, LLC, the OpenSearch Project, or the OpenSearch Community is implied by the use of these marks.

Python and its logo are registered trademarks of the Python Software Foundation in the United States and other countries. No endorsement by the Python Software Foundation is implied by the use of these marks.

React and its logo are registered trademarks of Meta Platforms, Inc. in the United States and other countries. No endorsement by Meta Platforms, Inc. is implied by the use of these marks.

Redis and its logo are registered trademarks of Redis Ltd. in the United States and other countries. No endorsement by Redis Ltd. is implied by the use of these marks.

Written by

Austin Greco

Staff Software Engineer

Austin Greco is a Staff Software Engineer on Michelangelo, Uber’s AI platform, where he led the developer experience and helped shape it from the early days into the system powering machine learning across the company.

Meghana Somasundara

Product Lead

Meghana Somasundara is a Product Lead for the Uber AI Platform. Previously she worked on building AI Platforms with other companies.

Frank Tenente

Sr. Legal Director, Product Legal (Platform)

Frank Tenente is Senior Legal Director, Mobility & Platform Product Legal at Uber, leading legal teams supporting Mobility and cross-platform business units. A commercial and product lawyer, Frank has scaled B2B legal teams and champions GenAI to transform contract review and legal work.

Sean Po

Staff Software Engineer

Sean Po is Staff Software Engineer and the Technical Lead on the Uber Agent Platform team.

Rush Tehrani

Sr Manager, Engineering

Rush is an Engineering Manager on the AI Platform Team at Uber. He supports the teams responsible for deployment and serving of classical, deep learning, and generative AI models, machine learning on mobile, and generative AI API gateway.

Related articles
4 articles
AI / ML
Backend
Engineering
October 1, 2026
AI / ML
Data
Engineering
September 22, 2026