diff --git a/README.md b/README.md index 6497f03c..15a5733d 100644 --- a/README.md +++ b/README.md @@ -21,6 +21,26 @@ more complex codebases. > Do not install Serena via an MCP or plugin marketplace! They contain outdated and suboptimal installation commands. > Instead, follow our [Quick Start](quick-start) instructions. +## What Our "End Users" Say + +Our end users are essentially AI agents, so they are also in the best position to evaluate Serena. +We wrote an unbiased and detailed evaluation prompt, which leads the agent to estimate the value-add of Serena's tools +when added on top of their existing capabilities. After about 20 minutes of thorough evaluation and experimentation, here +is what the agents had to say: + +**Opus 4.6 in Claude Code, evaluated on a large Python codebase:** +> "If I could ask my owner for one upgrade to my coding toolkit, it would be Serena — its symbolic addressing eliminated the constant re-reading that quietly eats my context on multi-edit sessions, its atomic cross-file refactorings collapsed 10-call manual workflows into one, and after a rigorous hands-on evaluation I can say that while built-in tools remain the right choice for small edits and text searches, Serena fills the exact gaps that make me feel clumsy without it." + +**Gpt 5.4 max in Codex CLI, evaluated on a Java codebase:** +> "As a coding AI agent, I would ask my owner to add Serena because it lets me work with code as stable symbols instead of fragile text, turning the +refactors, lookups, and multi-file edits that usually feel risky and repetitive into precise IDE-backed operations I can trust." + +See our documentation on the [evaluation methodology](https://oraios.github.io/serena/04-evaluation/000_intro.html) and the +detailed results beyond the brief recommendations above. +You can easily use our methods to run your own evaluation of Serena on a project of your choice. + +Your agent deserves the best coding tools, give them Serena! + ## How Serena Works Serena provides the necessary [tools](https://oraios.github.io/serena/01-about/035_tools.html) for coding workflows, @@ -224,4 +244,3 @@ Please refer to the [user guide](https://oraios.github.io/serena/02-usage/000_in A significant part of Serena, especially support for various languages, was contributed by the open source community. We are very grateful for the many contributors who made this possible and who played an important role in making Serena what it is today. - diff --git a/docs/04-evaluation/000_intro.md b/docs/04-evaluation/000_intro.md new file mode 100644 index 00000000..a66a8a9c --- /dev/null +++ b/docs/04-evaluation/000_intro.md @@ -0,0 +1,3 @@ +# Evaluation + +In this section we describe how we evaluate the performance of Serena's tools. diff --git a/docs/04-evaluation/010_evaluation-prompt.md b/docs/04-evaluation/010_evaluation-prompt.md new file mode 100644 index 00000000..540350e6 --- /dev/null +++ b/docs/04-evaluation/010_evaluation-prompt.md @@ -0,0 +1,149 @@ +# Evaluation Prompt + +Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of your choice. +All evaluations that you find in our documentation were created in one-shot sessions, only using this prompt and +then following up with a separate [summary prompt](011_followup-summary-prompt) + + +# Evaluate Serena's Tools Against Built-Ins + +You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both toolsets are used correctly. + +This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent user of both toolsets had only the built-ins, what concrete capabilities and efficiency wins would they be missing, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena adds — in terms of capabilities that weren't available at all, workflows that collapsed from many calls to one, and efficiency multipliers that show up across a session. Not a thumbs-up/thumbs-down, but a sharp description of the delta. + +Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope. They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add. + +**Describe the added value sharply, don't hedge.** A "both have their place" or "complementary, not rivals" opener is not a description of added value — it conveys zero information about what a reader would actually gain or lose by adding the tool. If Serena adds substantial capabilities, name them and quantify them. If it adds marginal capabilities, say that and show why. The two toolsets are complementary — that's a given, not the answer. The answer is a specific list of what Serena contributes to a correct-use workflow that built-ins alone cannot provide. + +## Ground rules + +### Starting conditions + +- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read documentation files either. Explore as if you've never seen it, focusing on code. +- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- ` or `git stash`. Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment. +- After each experiment, verify the working tree is clean with `git status --short` before moving on. + +### How to compare — correct use only + +- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse it. Examples of what *not* to report: +- Destructive write tools accepting what you send them — that is the contract of a destructive write, not a silent failure mode. +- Addressing schemes that require you to know what you're addressing — that is how addressing works. +- Tools refusing out-of-scope input (semantic tools on non-code files, single-expression refactorings on multi-branch functions, safe-delete on symbols with usages) — those are correct refusals. +- Mid-session friction that only appears if you mix tool families incorrectly — a caller managing their session properly doesn't hit it. +- Transient glitches may appear after `git checkout --` or other file-system mutations. Wait a moment before continuing after doing such a checkout, and in case of failure, retry once. If it then succeeds, the finding is a non-finding. +- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If you expect an error or "not applicable," don't make the call. +- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a broken input. + +### How to compare — workflow level, not single-call level + +- For every task, write out the full end-to-end call chain on each side before declaring a winner. Not just "Serena does it in one call" vs "Grep returns line N" — spell out every call you'd make to reach the goal, including the next step after whichever tool you called first. Many apparent wins and losses evaporate once you include the follow-up. +- Don't score a tool on criteria that only matter in the other toolset's workflow. If you find yourself penalizing Serena for missing a feature (line numbers, text anchors) or penalizing built-ins for missing a feature (name paths, type hierarchy), check whether that feature is actually needed in the tool's own native follow-up. If it's only needed because you're planning to fall back to the other toolset, you're holding the tool to the wrong workflow. +- Ephemeral addressing is a liability. Line numbers and byte offsets go stale the moment a file is edited. Stable addressing (name paths) is an efficiency win, even when the output looks smaller. + +### How to measure + +- Track observations while you work, not at the end. For every tool call, note: number of calls needed, approximate size of what you sent (including re-sent content), size of what you got back, and any prerequisite reads or follow-up verification. These are the raw material for the final report. +- Separate call count, input payload, output payload, and verification cost as distinct axes. A tool that halves call count but doubles payload may not be a net win. +- When comparing two approaches on the same task, include prerequisite Reads and post-hoc verification Greps in the cost — don't hide them outside the comparison. A cross-file rename via Edit isn't "one Edit call"; it's `grep + read × N + edit × N + verify`. + +## Exploration phase — tasks to actually perform + +Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an item isn't applicable ("no suitable candidate in this codebase" is a valid reason). + +### Codebase understanding + +1. Get a high-level overview of the repository structure — top-level layout, main packages, entry points. +2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with Glob/Grep/Read. Then write out the concrete next step on each side ("after this overview, to read the body of method X I'd call _____") and compare the pair of calls, not just the overview call. Whichever tool's output most directly feeds its own follow-up wins on workflow terms. +3. Pick a specific method inside a class and retrieve its body without reading the surrounding file. +4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the question "who uses this in code?" vs "where is this mentioned anywhere, including docs?" — these are different questions, and each toolset is naturally suited to one of them. +5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what text search would need to do (and whether it can follow an override chain or cross into stub files in one step). +6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or signature. Note whether each toolset can do this at all and what infrastructure it requires (env activation, site-packages discovery, etc.). + +### Single-file edits — span the full range of edit sizes + +Do all three sizes below, not just one. The comparison between content-anchored editing (Edit) and symbol-body replacement is size-dependent: Edit's payload grows with the old+new anchor pair, while symbolic body-replace grows with the full new body. They cross over, and where they cross is the whole point. Only testing small edits hides the crossover. + +7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method. Do it with `Edit` and with `replace_symbol_body`. Compare payload sent, payload received, and prerequisite reads. + +7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its signature. Do it both ways. Note that for Edit you may need a long anchor to keep the old_string unique, and that for symbolic replacement the new body is similar in size to what you'd send for Edit's new_string alone. + +7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways. This is where symbolic body replacement is designed to win: Edit has to send the entire old body as an anchor and the entire new body, while `replace_symbol_body` sends only the new body + a short name path. Measure the ratio. + +8. Insert a new function/method at a specific structural location (e.g., "right after this existing method"). Try both the symbolic-insert path and the manual Edit path. +9. Rename a private helper used only within one file. Compare doing it by hand vs. using a semantic rename. + +### Multi-file changes + +10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path against the built-in equivalent chain. Under correct use, a semantic rename is paired with a short post-rename `Grep` for text-surface references (docstrings, markdown, notebooks) — count that as a complementary step, not a Serena failure. +11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if available; plan the built-in equivalent honestly (how many Reads, Edits, and import-cleanup decisions would it take?). +12. Delete a symbol safely, checking it has no remaining usages. Compare "search-then-delete" with a safe-delete tool. +13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable (single-expression body, no early returns, no side effects, substitutable at its call sites). If no such candidate exists, report "no suitable candidate" and skip. Do not contrive a broken input. + +### Reliability & correctness under correct use + +14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class method, override, or overload that text search would over-match. The point is to show the capability — precision under correct use — not to manufacture a rename that breaks polymorphism through careless targeting. +15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit` calls is not. You don't have to force a failure — report on what this means for reliability on real failures (disk full, interrupted process, transient permission errors). +16. Success signals. For each completed refactor, note what each tool returns on success. Both toolsets report mechanical success only; semantic intent is always the caller's responsibility to verify with a diff. Note this as a baseline for both sides, not as a weakness of either. + +### Workflow effects across multiple edits + +17. Chain at least three edits in one file. Report what each toolset requires between edits. Pay particular attention to whether identifiers survive mutation: name paths stay valid across edits in other regions of the file; line numbers and byte offsets don't. This is the single biggest workflow-level efficiency effect. +18. Multi-step exploration across the repo. Note whether intermediate results (overviews, reference lists) remain useful across later edits, or whether they have to be refreshed. Stable output survives a session; ephemeral output does not. + +### Things where the comparison shouldn't be interesting + +19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use `Read`. State the applicability boundary once and move on. +20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`. Don't call symbolic search on free text. + +## Evaluation phase + +Write a report structured for progressive disclosure — lead with the strongest insights; a reader should be able to stop at any point and still walk away informed. + +**Value-weighting is not optional.** For every contribution you name — in §1, §2, §7, and anywhere else you list what Serena adds — you must estimate *how much it matters in general coding work*, not just whether it's novel. A capability that saves 10 calls but is used once a month is a smaller practical contribution than a capability that saves 1 call but is used every editing session, and a reader needs to be able to tell which kind of contribution each item is. Be explicit about **frequency** (how often does this matter in typical Python coding?) and **value per hit** (when it matters, how much does it save?). Order by the product, not by how impressive the individual feature sounds. + +**Every section must end with a one-sentence verdict** (a short paragraph labelled "**Verdict:** ...") that gives the reader the single-sentence takeaway for that section. This applies to every top-level section §1–§9 and to each subsection under §3. The verdict is a recommendation in context — e.g. "use symbolic addressing whenever you expect multiple edits to one file," "every task in this group is a clear Serena win," "built-ins only; this is not a contest." Not a hedge, not a summary — a pointed one-liner a reader can act on. + +1. **Headline: what Serena adds.** Open with a sharp description of the delta Serena provides on top of the built-ins — not a thumbs-up/thumbs-down, not a "they're complementary" hedge, but a specific list of what Serena contributes that built-ins alone cannot. Structure it as a short list of *capabilities* (things that become possible) and *efficiency multipliers* (things that get cheaper by how much), **ordered by how much value each contribution actually delivers in general coding work — not by novelty**. For each item, state frequency (how often does this matter?) and value per hit (when it matters, how much?). Group items into high/medium/low value tiers if the gap is large enough that a reader should care about the ordering. A reader stopping after this section should know not only what the added value is but also how much of their daily work it touches. If your opening paragraph could be written about a toolset that adds *nothing*, rewrite it. End with a one-sentence verdict. +2. **Added value, by area (3–6 bullets).** Each bullet answers: *what would a built-ins-only workflow be missing here, and how much would it miss it?* Lead with the areas of largest weighted value — not the most novel capability, but the one whose frequency × value-per-hit product is biggest. Each bullet must include a concrete frequency estimate and a concrete value-per-hit estimate (in calls saved, tokens saved, or correctness improved). Bullets should describe what Serena contributes, not "wins" or "losses" against the built-ins. If you catch yourself writing "Serena wins at X," rewrite as "Serena adds X, which shows up in [frequency] coding work and saves [value]." End with a one-sentence verdict. +3. **Detailed evidence, grouped by capability.** Per-task: what you tried, the full end-to-end call chain on each side (including prerequisite reads and complementary follow-up steps), payloads sent and received. Be specific — "1 call vs ~10, and the Serena call sent ~200 tokens while the Edit equivalent would have sent ~450 after including the prerequisite Read" is useful; "Serena was faster" is not. End each subsection with its own one-sentence verdict in context (e.g. "Verdict (multi-file refactors): every task in this group is a clear Serena win"). +4. **Token-efficiency analysis.** Separate from raw call count. Address: + - Payload asymmetry as a function of edit size. Where's the crossover between content-anchored editing and symbolic body replacement? Show the ratio at small, medium, and large edit sizes. + - Forced reads: does a tool make you load content into context that you don't actually need? + - Stable vs ephemeral addressing. Name paths survive edits; line numbers don't. The output-size comparison must account for shelf life, not just raw tokens returned. + + End with a one-sentence verdict on when token economics favor each toolset. +5. **Reliability & correctness analysis** — under correct use. Address: + - Precision of matching: semantic identifiers vs textual matches, with concrete cases where the question shape determined which tool was the right one. + - Scope disambiguation across override chains and overloads. + - Atomicity of multi-file operations, and what it means on real failures. + - Transitive semantic queries (type hierarchy, reference chains, dependency lookups) that text search cannot approximate in one step. + + End with a one-sentence verdict on where correctness weight favors each toolset. +6. **Workflow effects across a multi-step session.** How do efficiency gaps widen over a session? Do identifiers survive edits? Does intermediate output (overviews, reference lists) stay useful? This is often where the real gap lives — a one-call comparison can miss it. End with a one-sentence verdict. +7. **Capabilities with no built-in equivalent.** What Serena made possible that the built-ins genuinely cannot do, or can only approximate with a much longer workflow. Name each capability individually **and annotate each with how much it matters in general coding work** — a capability that's unique but rarely needed is a smaller contribution than one that's unique and needed constantly. End with a one-sentence verdict. +8. **Where built-ins remain the right default.** Which tasks are still best served by Grep/Read/Edit, and why. Include the non-code-file boundary and the post-rename text sweep explicitly. Estimate what share of daily coding these cases represent — this is what calibrates §1's weighting. End with a one-sentence verdict. +9. **Usage rule for a developer with both toolsets.** A per-task decision rule: given both toolsets installed, which do you reach for and when? This is the practical takeaway, not the headline — the description of added value belongs in §1. End with a one-sentence verdict. + +## What I'm looking for + +Ground every claim in something you actually measured or observed. If an initial impression turns out to be wrong as you gather more evidence, update the report. + +**Pay particular attention to the "what does it add, and how much" question.** The §1 headline is the part a reader will actually remember. Before you write it, ask yourself: if a developer reads only that section, will they know (a) what specific capabilities and efficiency wins Serena contributes, and (b) how much those contributions actually matter in general coding work? A reader should be able to tell the difference between "rare but large" contributions and "common but small" ones without guessing. If the answer is no, rewrite it. + +Also pay attention to: + +- **Weighted value, not raw novelty.** A capability that saves 10 calls but is used once a month is a smaller practical contribution than one that saves 1 call but shows up every editing session. Your §1 and §2 ordering must reflect frequency × value-per-hit, not how impressive the individual feature sounds. Name both axes explicitly for each contribution. +- **Second-order effects visible only across a session.** Token cost of re-sending content, whether identifiers survive mutation, atomicity on failure, intermediate output staying valid. A one-call comparison misses these; a multi-edit session exposes them. These are often high-frequency contributions that look small per hit. +- **Workflow-level honesty.** Compare the full call chain on each side, not just the first call. Don't let one toolset's vocabulary set the evaluation axes. +- **Capability deltas under correct use.** What does Serena let you do that you genuinely couldn't do — or couldn't do cheaply — with built-ins alone? These are the findings that justify the heaviest weights in §1. + +## What I am not looking for + +- **Failures from misuse.** Calling a code-semantic tool on a non-code file, inlining an un-inlineable function, renaming without a valid name path, passing a body that drops a branch to `replace_symbol_body` without first reading the current body. These are caller errors, not tool findings. +- **Gotcha hunts and "symmetric failure mode" comparisons.** If you find yourself writing "tool X silently does Y when misused" or "tool A has a loud failure that tool B lacks," stop. You've drifted from evaluation into user guide. A destructive write accepting what you send it is a contract, not a flaw; an addressing scheme that forces you to see what you're addressing is a consequence of addressing, not a safety feature. Report capability and efficiency, not caller-discipline requirements. +- **Mid-session friction from tool mixing.** If a workflow only breaks because you're switching tool families mid-edit on the same file, that's a usage choice, not a tool finding. +- **Transient glitches after `git checkout --` or other file-system mutations.** Retry once; if it then succeeds, don't include it. +- **Hedged or neutral §1 openings.** "Both toolsets have their place" and "they're complementary, not rivals" are information-free openings that could be written about any tool pair. §1 must describe Serena's specific contribution — capabilities and efficiency multipliers — in a way that would not be true of a toolset that adds nothing. If your opening paragraph could be recycled for another tool, rewrite it. +- **Binary adoption pitches.** "Install it" and "don't bother" are also the wrong shape of answer. The question isn't whether to adopt — it's what the added value actually is, and how much of it. A good §1 lets the reader decide for themselves whether the described delta is worth it for their work, by being specific about what the delta is and how often it shows up. +- **Novelty-weighted ordering.** Listing Serena's most exotic capability first because it's the most impressive one misleads a reader about what they'll actually experience. If the most exotic feature is used once a month and a boring one is used every session, the boring one comes first. A contribution's weight is frequency × value-per-hit, not how surprising it is. +- **Unquantified "wins."** "Serena is faster here," "Serena is more reliable," "built-ins are cheaper" are all unhelpful without a quantity attached. Every claim about relative value needs a concrete number or a concrete frequency — calls saved, tokens saved, times per session, share of daily edits. An evaluation without magnitudes is a review, not a measurement. diff --git a/docs/04-evaluation/011_followup-summary-prompt.md b/docs/04-evaluation/011_followup-summary-prompt.md new file mode 100644 index 00000000..cb5ac6e0 --- /dev/null +++ b/docs/04-evaluation/011_followup-summary-prompt.md @@ -0,0 +1,9 @@ +# Post-Evaluation Summary Prompt + +_Use the prompt below to summarize the evaluation of Serena by your agent. +We used it to create the recommendations we display in the readme._ + +Read the evaluation reperot and write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit +of emotion but grounded in the evaluation that you have done above. It should be written from the perspective of a coding AI agent and make it +clear whether an AI agent would ask its owner to add Serena's tools or not. It will be the first sentence that a potential new user will see +about how useful Serena is for an AI agent. \ No newline at end of file diff --git a/docs/04-evaluation/results/000_intro.md b/docs/04-evaluation/results/000_intro.md new file mode 100644 index 00000000..c32fea83 --- /dev/null +++ b/docs/04-evaluation/results/000_intro.md @@ -0,0 +1,3 @@ +# Results + +Here are the results of the evaluation for various coding agents running on various projects. diff --git a/docs/04-evaluation/results/010_cc_on_tianshou-serena-evaluation.md b/docs/04-evaluation/results/010_cc_on_tianshou-serena-evaluation.md new file mode 100644 index 00000000..e6a31854 --- /dev/null +++ b/docs/04-evaluation/results/010_cc_on_tianshou-serena-evaluation.md @@ -0,0 +1,339 @@ +# Evaluation by Opus 4.6 in Claude Code + +## Final Recommendation + +The recommendation is created with the [follow-up summary prompt](../011_followup-summary-prompt), everything else was created in one shot from the +[evaluation prompt](../010_evaluation-prompt). + +> "If I could ask my owner for one upgrade to my coding toolkit, it would be Serena — its symbolic addressing eliminated the constant re-reading +that quietly eats my context on multi-edit sessions, its atomic cross-file refactorings collapsed 10-call manual workflows into one, and after a +rigorous hands-on evaluation I can say that while built-in tools remain the right choice for small edits and text searches, Serena fills the exact +gaps that make me feel clumsy without it." + + +Ground-truth hands-on evaluation run on `tianshou` (Python, ~26 kLOC under `tianshou/`), focused on what Serena's tools add on top of `Read/Edit/Write/Glob/Grep` under correct use. Every measurement below comes from an actual tool call on this repo, not a thought experiment; edits were applied and reverted via `git checkout --` between experiments. + +--- + +## §1. Headline — what Serena adds + +On a typical Python editing session, Serena contributes four things built-ins cannot reproduce cheaply or at all. Ranked by **frequency × value-per-hit**, highest first: + +**Tier A — touches almost every editing session** + +1. **Stable name-path addressing that survives intermediate edits.** Every Serena editing and querying tool is addressed by `ClassName/method_name`, which remains valid after inserts, deletes, and rewrites elsewhere in the same file. Grep/Read/Edit chain through line numbers or text anchors, both of which drift. **Frequency: every multi-edit session, which is most of them.** **Value per hit: saves one Read + one re-grep per subsequent edit in the same file (~400–800 bytes of re-sent content each time).** This is the quietest contribution and probably the largest in aggregate. + +2. **Payload asymmetry on medium-and-larger body rewrites.** `replace_symbol_body` sends `name_path + new_body`, while `Edit` sends `old_body + new_body`. Measured crossover is around 10–15 lines: below that, Edit's tiny anchor wins 5–10× on payload; above that, Serena wins ~2× at 20 lines and ~2.3× at 66 lines, and the gap keeps widening with body size. **Frequency: once per session you rewrite a non-trivial method body.** **Value per hit: 50% payload reduction on medium rewrites, climbing toward ~60% on large ones, plus no need for a prior Read to capture the exact anchor.** + +**Tier B — rare per session but very high value when it fires** + +3. **Single-call cross-file refactorings with atomicity and automatic import cleanup.** `rename`, `move`, and `safe_delete` each collapse an 8–12-call manual workflow (grep → read-all-callers → edit each → verify → hand-clean unused imports) into one call, atomically. **Frequency: 0–3 times per session, depending on task type — zero on many days, constant on refactoring days.** **Value per hit: 5–10× call reduction plus correctness improvements that manual chains drop (I observed `move` auto-remove a `Categorical` import from the source file that a human refactorer would easily miss).** + +4. **Semantic navigation into third-party dependency source.** `find_symbol(search_deps=True)` / `find_declaration` resolves a type from its use site straight into `site-packages` or stub files, returning the full class body in one call. Built-ins would require: (a) reading the file to find the import, (b) locating the venv, (c) guessing the module path, (d) reading the file. **Frequency: a few times per session when debugging type errors or reviewing unfamiliar library APIs.** **Value per hit: 3–4 manual calls and a path-guessing step collapsed into one.** + +**Tier C — capability-level, not efficiency-level** + +5. **Transitive type hierarchy and reference graphs.** `type_hierarchy` returns the full super/sub chain — including into external stub files (`.pyi`) — in one call. `find_referencing_symbols` returns each reference *with its containing symbol* (the function/class that holds it), which Grep cannot produce. Built-ins can imitate single-level relationships with Grep but cannot chase override chains in one step. **Frequency: once or twice a session on unfamiliar code; close to never once you know the hierarchy.** **Value per hit: one call vs. N iterative Greps, plus recall into stub files.** + +**Verdict:** Serena's biggest practical contribution is the *quietest* one — stable symbolic addressing that survives in-file mutations — and its most *spectacular* one — single-call atomic cross-file refactorings — is rare but decisive when it fires; the middle-weight wins are payload asymmetry on body rewrites and cheap dependency navigation. + +--- + +## §2. Added value, by area + +Ordered by frequency × value-per-hit, not novelty. + +- **Stable identifiers across chained edits to one file.** Demonstrated in Task 17: three sequential edits to `CollectStats` (`insert_after refresh_len_stats`, `insert_after refresh_return_stats`, `replace_symbol_body refresh_std_array_stats`) all used the same kind of stable name-path address and worked without re-reading the file. An equivalent `Edit` chain after the first two inserts would have needed at least one intermediate `Read` because line numbers and some text anchors shift. **Frequency: every multi-edit session.** **Value: ~1 re-read (~400–800 bytes) saved per subsequent edit; compounds across a session.** + +- **Body replacement at medium/large sizes.** Measured in Tasks 7a/7b/7c on `collector.py`: + - 1-line change (7a): Edit `~120 B`, `replace_symbol_body` `~1200 B`. **Edit wins ~10×.** + - ~20-line rewrite (7b): Edit `~1430 B`, Serena `~715 B`. **Serena wins ~2×.** + - ~66-line rewrite (7c): Edit `~4700 B`, Serena `~2050 B`. **Serena wins ~2.3×.** + Crossover sits around 10–15 lines. **Frequency: once or twice per substantive editing session.** **Value: 50–60% payload cut on the edits where it applies; grows with body size.** + +- **Atomic cross-file operations.** Task 10 renamed `EpisodeRolloutHookMCReturn` across `collector.py` + `test_collector.py` in one `rename` call, and — importantly — *also* updated a Sphinx `:class:` cross-reference in a docstring that a pure `Edit` chain would only catch if the caller remembered to sweep docstrings separately. Task 11 moved `get_stddev_from_dist` from `collector.py` to `batch.py`: one call updated the target file, added the import at the source, updated a separate test file's import, *and* removed the now-unused `Categorical` import from the source. **Frequency: 0–3 times per session.** **Value per hit: 5–10 calls saved plus one or two automatic correctness wins the manual path tends to miss.** + +- **Dependency navigation.** Task 6 resolved `Distribution` from its use site in `collector.py` into the real `torch.distributions.distribution.Distribution` class body (~300 lines) with a single `find_declaration` call. The built-in path requires parsing imports, guessing the venv site-packages location, and reading the file by hand. **Frequency: a few times per session on unfamiliar code or type debugging.** **Value: 3–4 calls plus one guess collapsed to one.** + +- **Reference and hierarchy queries that distinguish "code usage" from "any text match".** Task 4 on `Collector`: `find_referencing_symbols` returned ~70 files with each reference tagged by containing symbol (`Py:IMPORT_ELEMENT`, `Py:FUNCTION_DECLARATION: ["test_dqn"]`), while `Grep \bCollector\b` returned 332 matches across 88 files including notebooks, SVGs, `CHANGELOG.md`, and `README.md`. These are answers to *different questions*, and each tool is naturally matched to one. **Frequency: several times per session; the right question determines the right tool.** **Value: precision saves a read of each false-positive file (~10–30 saved reads on a popular class).** + +- **Safe-delete with automatic usage check.** Task 12: `safe_delete` on `EpisodeRolloutHookMerged` succeeded silently (no usages); on `CollectStats` it refused with a 200+ line usage list grouped by enclosing symbol. The built-in equivalent is a `Grep` followed by manual discipline — same call count but no enforcement. **Frequency: rare (delete-by-name operations).** **Value: when it fires, it prevents the most expensive class of mistake (deleting a referenced symbol).** + +**Verdict:** The two highest-weighted contributions are the boring one (stable in-file addressing, every session) and the refactoring one (rare but 5–10× call reduction); Serena's more exotic capabilities (hierarchy, dependency navigation) are lower-frequency polish on top. + +--- + +## §3. Detailed evidence, grouped by capability + +### §3.1 File structural overview + +**Task 2** — structural overview of `tianshou/data/collector.py` (1551 lines). +- Serena `get_symbols_overview(depth=0)`: ~240 bytes returned; top-level classes + functions names only, no line numbers, no per-class method detail. +- Serena `get_symbols_overview(depth=1)`: ~3200 bytes; classes with nested method and field lists. +- Grep `^(class |def | def )`: ~3300 bytes; identical structural information plus line numbers. +- Output size is roughly equivalent at `depth=1`. The real difference shows up in the *follow-up* call: + - Serena follow-up (`find_symbol("Collector/_compute_action_policy_hidden", include_body=True)`): 1 call, returns body only, ~66 lines. Address `Collector/_compute_action_policy_hidden` remains stable across any edits to unrelated regions. + - Built-in follow-up (`Read offset=707 limit=66`): 1 call, returns body + line-number prefixes. Address (line 707) goes stale after any insert above that line. +- The winner depends on what comes next: if you stop at reading, they tie; if you plan to edit the body next, Serena's address is the one that survives the subsequent mutation. + +**Verdict (§3.1):** tie on the overview call alone; Serena wins on the full read-then-edit chain because of address stability. + +### §3.2 Method body retrieval + +**Task 3** — body of `Collector._compute_action_policy_hidden`. +- Serena: 1 call (`find_symbol include_body`), ~1800 bytes of body-only output. +- Built-ins: 1 call (`Read offset=707 limit=66`), ~1900 bytes including line prefixes. +- Essentially equivalent on a cold cache; Serena's output is cleaner for machine consumption, built-ins' is cleaner for humans reviewing line locations. + +**Verdict (§3.2):** tie in isolation; Serena wins when followed by an edit. + +### §3.3 Reference search + +**Task 4** — who uses `Collector`? +- Serena `find_referencing_symbols`: ~70 files, each reference annotated with its containing symbol. Precise to code usage. Missed docstring/README mentions almost entirely (except a few `FILE`-tagged entries). +- Grep `\bCollector\b`: 332 occurrences across 88 files, includes `docs/*.ipynb`, `docs/*.md`, `structure.svg`, `CHANGELOG.md`, `README.md`. +- Cost comparison for the question "every code file that *uses* this class": + - Serena: 1 call, answer usable directly. + - Grep: 1 call, answer needs filtering by extension and visual dedup. +- Cost comparison for "anywhere this name appears in the repo, including docs": + - Serena: cannot answer directly. + - Grep: 1 call, answer usable directly. + +**Verdict (§3.3):** each toolset wins on its own question; pick by question shape, and do not penalize either for missing the other's answer. + +### §3.4 Type hierarchy and override chains + +**Task 5** — super/sub types of `BaseCollector`. +- Serena `type_hierarchy(depth=0)`: 1 call, returns transitive hierarchy including `Collector → AsyncCollector` on the sub side and `ABC → object` on the super side — the super side is resolved into an external `` stub file. +- Grep `class \w+\(.*BaseCollector`: 1 call, returns exactly one hit (`Collector`). To find `AsyncCollector` I'd have to issue a second grep for `class \w+\(.*Collector`, and so on recursively. External supertypes cannot be reached at all. +- **This is a capability delta, not a raw efficiency delta**: text search cannot cross an override chain in one step, and cannot reach `.pyi` stubs. + +**Verdict (§3.4):** clear Serena capability win; value grows with hierarchy depth and cross-file reach. + +### §3.5 Dependency navigation + +**Task 6** — resolve `Distribution` from its usage in `collector.py`. +- Serena `find_declaration` with a regex anchored at the usage site: 1 call, returned the full 300-line body of `torch/distributions/distribution.py::Distribution`, including docstrings. Unambiguous — the tool used the import context to pick the right file out of 41 candidates named `Distribution` in the dependency tree. +- Built-in equivalent: (a) Read `collector.py` imports, (b) map `from torch.distributions import Distribution` to `torch/distributions/distribution.py`, (c) locate the venv's site-packages path, (d) Read the file. 3–4 calls plus one implicit "where is the venv" step. + +**Verdict (§3.5):** clear Serena capability win whenever you need third-party source. + +### §3.6 Small edits (< ~10 lines) + +**Task 7a** — one-line error message change inside `BaseCollector._validate_buffer` (21-line body). +- `Edit old_string=".. should be greater than 0." new_string=".. must be strictly positive."`: ~120 bytes on the wire, one call. Prerequisite: know the exact anchor (one prior Grep or Read, which I already had from earlier in the session). +- `replace_symbol_body`: ~1200 bytes (whole 21-line body resent). Prerequisite: prior `find_symbol include_body` to see the current body. +- Result diff identical in both cases (`1 insertion(+), 1 deletion(-)`). + +**Verdict (§3.6):** Edit wins ~10× on payload for ≤ 3-line changes; use Edit whenever the change is small and the anchor is obvious. + +### §3.7 Medium edits (~10–30 lines) + +**Task 7b** — rewrite ~20 lines inside `CollectStats.update_at_step_batch`. +- Edit: ~750 B old anchor + ~680 B new body = ~1430 B on the wire. +- `replace_symbol_body`: ~680 B new body + ~35 B name_path = ~715 B. +- Both end with identical diffs. +- Prerequisite is symmetric: both need to see the current body first (one `find_symbol` or one `Read`). + +**Verdict (§3.7):** crossover point — Serena pulls ahead around 15 lines and wins ~2× at 20. + +### §3.8 Large edits (50+ lines) + +**Task 7c** — whole-body rewrite of `Collector._compute_action_policy_hidden` (~66 lines). +- Edit: ~2700 B old anchor + ~2000 B new body = ~4700 B. +- `replace_symbol_body`: ~2000 B new body + name_path = ~2050 B. +- Ratio ~2.3×. The asymmetry grows linearly with body size: Edit scales as `O(old + new)`, `replace_symbol_body` as `O(new)`. + +**Verdict (§3.8):** clear Serena win; reach for `replace_symbol_body` on any substantial method-body rewrite. + +### §3.9 Structural insertion + +**Task 8** — insert a new method right after `CollectStats.refresh_len_stats`. +- Serena `insert_after_symbol` with `name_path="CollectStats/refresh_len_stats"` and body of the new method: ~150 B on the wire, unambiguous anchor, 1 call. Result: new method inserted cleanly between `refresh_len_stats` and `refresh_std_array_stats`. +- Edit equivalent: old_string must capture the end of `refresh_len_stats` uniquely — roughly the last 5 lines of the method — then new_string replays those lines and appends the new method. Approximately 550 B on the wire. + +**Verdict (§3.9):** Serena wins ~3× on payload and on anchor stability for structural inserts. + +### §3.10 Single-file rename of a private helper + +**Task 9** — rename `_HACKY_create_info_batch` → `_create_info_batch_legacy` (2 occurrences, same file). +- Serena `rename` via name_path: 1 call. +- Edit with `replace_all=true`: 1 call. Both the declaration and the one call site get rewritten. +- Both succeed; both leave a clean 2-line diff. + +**Verdict (§3.10):** tie — for single-file renames of a distinctive identifier, built-in `Edit(replace_all=true)` is competitive. + +### §3.11 Multi-file rename + +**Task 10** — rename `EpisodeRolloutHookMCReturn` → `EpisodeRolloutMCReturnHook` (5 sites across `tianshou/data/collector.py` and `test/base/test_collector.py`). +- Serena `rename`: 1 call, atomic across both files. Also updated a Sphinx `:class:` docstring cross-reference at `collector.py:611` via the default `rename_in_comments=True`. +- Built-in chain: `Grep` (1) → `Edit replace_all=true` on `collector.py` (1) → `Edit replace_all=true` on `test_collector.py` (1) → verification `Grep` (1). 4 calls minimum, 5–6 if I remember to hit the docstring reference. Not atomic across files. + +**Verdict (§3.11):** Serena wins ~4–5× on call count and catches the docstring reference by default. + +### §3.12 Cross-module move + +**Task 11** — move `get_stddev_from_dist` from `tianshou/data/collector.py` to `tianshou/data/batch.py`. +- Serena `move`: 1 call. Final `git diff --stat`: + ``` + test/base/test_stats.py | 3 ++- + tianshou/data/batch.py | 24 ++++++++++++++++++++++++ + tianshou/data/collector.py | 27 ++------------------------- + ``` + The tool: + 1. Inserted the function into `batch.py`. + 2. Removed it from `collector.py`. + 3. Added `from tianshou.data.batch import get_stddev_from_dist` to `collector.py`. + 4. Updated `test/base/test_stats.py`'s import of this function. + 5. **Removed the now-unused `Categorical` import from `collector.py`**, because nothing else in that file referenced it. +- Built-in chain (honestly planned): Grep for callers (1) → Read each caller to see current import form (≥ 3) → Edit collector.py to remove the function (1) → Edit batch.py to insert (1) → Edit each caller's import (≥ 2) → Edit collector.py to remove the now-unused `Categorical` import (1) → verification Grep (1). **Roughly 10 calls**, and the "unused import" cleanup is the kind of thing a distracted human misses. + +**Verdict (§3.12):** biggest call-count collapse in the evaluation — ~10× — with a correctness bonus. + +### §3.13 Safe deletion + +**Task 12** — two deletions. +- `safe_delete(EpisodeRolloutHookMerged)` (unused): succeeded, removed the class cleanly, 38-line deletion. +- `safe_delete(CollectStats)` (heavily used): refused with an explicit `SafeDeleteFailedException`, returning ~220 usage locations grouped by enclosing symbol. +- Built-in equivalent: `Grep` for usages → visual inspection → `Edit` deletion. Same call count when the symbol is unused; similar when it isn't, but with no enforcement — a careless `Edit` would happily delete a used symbol and break the repo. + +**Verdict (§3.13):** same call count as manual, but the enforcement eliminates the worst mistake class. + +### §3.14 Inline helper + +**Task 13** — skipped. No legally inlinable helper (single-expression, side-effect-free, with call sites) found in a quick scan of `tianshou/`. Per the prompt, I did not contrive a broken input. + +**Verdict (§3.14):** no data; comparison not applicable on this codebase. + +### §3.15 Scope precision and disambiguation + +**Task 14** — find every `_collect` in `collector.py`. +- `find_symbol("_collect")` returned three distinct hits with their full name paths (`BaseCollector/_collect`, `Collector/_collect`, `AsyncCollector/_collect`), plus inlined signatures for each. Each is addressable unambiguously — I can rename, replace-body, or reference exactly one of them. +- `Grep "_collect\("` would return all three locations but with no structural distinction: to tell them apart I'd have to read surrounding context to see which class each `def` belongs to. + +**Verdict (§3.15):** Serena's addressing is precise by construction on override chains and overloads. + +### §3.16 Chained edits to one file + +**Task 17** — three successive edits on `CollectStats`: +1. `insert_after_symbol("CollectStats/refresh_len_stats", ...)` — insert new `reset_len_stats` method. +2. `insert_after_symbol("CollectStats/refresh_return_stats", ...)` — insert new `reset_return_stats` method. +3. `replace_symbol_body("CollectStats/refresh_std_array_stats", ...)` — rewrite an existing method body. + +All three calls used the original name paths, unchanged; none required a re-Read between edits. Final `git diff` showed exactly the expected three-edit composite: two insertions plus one body rewrite, adjacent and clean. The first two inserts shifted line numbers between 8 and 16 lines, which would have invalidated any line-number addressing for the third target. + +An equivalent Edit chain would have survived this particular sequence (because the anchors were text, not line numbers), but it exposes the *general* pattern: name_paths are mutation-proof addresses, line numbers are not, and large text anchors become non-unique quickly. + +**Verdict (§3.16):** for any session with three or more edits to the same file, symbolic addressing is a systematic safety and efficiency win. + +### §3.17 Non-code files and free-text searches + +**Tasks 19/20** — state the applicability boundary and move on. Semantic tools don't apply to changelogs, notebooks, configs, or free-text searches for log strings; `Read` and `Grep` are the right tools. + +**Verdict (§3.17):** built-ins only; not a contest. + +--- + +## §4. Token-efficiency analysis + +**Payload asymmetry by edit size** (measured on `collector.py`): + +| Edit size | Edit (old+new) | `replace_symbol_body` | Ratio | +|---|---|---|---| +| 1 line (in 21-line method) | ~120 B | ~1200 B | **Edit 10× smaller** | +| ~20 lines rewritten | ~1430 B | ~715 B | **Serena 2× smaller** | +| ~66 lines whole-body | ~4700 B | ~2050 B | **Serena 2.3× smaller** | + +Crossover is around 10–15 lines. Below it, Edit's per-change payload is dominated by the tiny anchor and wins by an order of magnitude. Above it, Edit pays once for the old body and again for the new, while `replace_symbol_body` pays only for the new body plus a ~30-character name_path; the gap grows linearly with body size. Structural inserts (`insert_after_symbol`) have similar asymmetry — the name_path replaces a multi-line text anchor. + +**Forced reads.** Serena's overview tools return symbol names without bodies and its reference tools return containing-symbol metadata without snippets, so you control when to pull code into context. `find_symbol(include_body=False)` + `find_symbol(include_body=True, name_path=...)` is a two-step "browse, then fetch body" pattern that keeps context lean; the built-in equivalent is `Grep` (which does not return bodies) + `Read` with an offset/limit, which is about as lean but requires the caller to compute the limit by hand. + +**Stable vs ephemeral addressing.** Name paths remain valid across edits to unrelated regions of the same file. Line numbers and byte offsets do not, and text anchors become ambiguous once a file grows. The output-size comparison has to account for *shelf life*: a slightly larger overview that stays useful across an entire session is cheaper than a slightly smaller one that has to be regenerated after each edit. In a five-edit session on one file, Serena's name-path overview is queried once; the line-number grep output is effectively refreshed after each edit that shifts upstream content — an O(edits × file_grep_cost) hidden tax that the one-call comparison misses. + +**Verdict (§4):** under ~10-line edits, built-in `Edit` is cheaper; above that threshold, symbolic body replacement wins on raw payload; across a multi-edit session, name-path addressing wins regardless of size because it doesn't decay. + +--- + +## §5. Reliability and correctness analysis (under correct use) + +**Precision of matching.** `find_referencing_symbols` on `Collector` returns ~70 code files and annotates each with the containing symbol that holds the reference. `Grep \bCollector\b` returns 332 hits across 88 files including notebooks, SVGs, and markdown. For the question "which Python files import and use this class?", Serena's output is directly usable and Grep's needs a filter pass. For the question "where is the name `Collector` mentioned anywhere in the repo, including docs, changelog, and diagrams?", Grep's output is directly usable and Serena's is incomplete by design. Each tool's precision is perfect *for its question*; the mistake is asking the wrong one. + +**Scope disambiguation across overrides.** `find_symbol("_collect")` returned three distinct name paths — `BaseCollector/_collect`, `Collector/_collect`, `AsyncCollector/_collect` — each independently addressable for rename or body replacement. Text search on `_collect(` returns three line locations with no structural annotation; the caller must read context to tell them apart, and any cross-file rename risks touching the wrong override if called carelessly. + +**Atomicity on real failures.** `move` of `get_stddev_from_dist` made coordinated changes across three files in one call. A five-step Edit chain replicating the same move would leave the repo in a half-renamed state if any intermediate step failed on disk-full, a permission error, or an interrupted process; a single-call atomic refactoring either completes or leaves the working tree clean. This matters less for local agent sessions (where the blast radius of a partial refactor is small and recoverable with `git checkout --`) and more for any workflow where a failed run is committed or pushed. + +**Transitive semantic queries.** `type_hierarchy` returned both the sub-chain (`Collector → AsyncCollector`) and the super-chain (`BaseCollector → ABC → object`, with `ABC` resolved into an external `.pyi` stub) in one call. No text-search workflow reaches into stub files or chains override relationships in one step. + +**Success signals (symmetric).** Both toolsets return only mechanical success: "the file was written," "the rename finished." Neither verifies that the new code still compiles, type-checks, or preserves semantics. Post-edit `git diff` review is the caller's responsibility on both sides. + +**Verdict (§5):** Serena's correctness edge is concentrated in questions that are semantic by nature — override chains, transitive type queries, cross-file atomic ops — and tied to Grep/Edit on questions that are textual by nature. + +--- + +## §6. Workflow effects across a multi-step session + +The single biggest session-level effect is **identifier stability across edits**. A session that makes 5 edits to `collector.py` looks like: + +- With symbolic addressing: one `get_symbols_overview` at the start, five `replace_symbol_body`/`insert_after_symbol` calls by name path. The overview is consulted zero or one more times. No re-reads of the file between edits. +- With line-number or large-text-anchor addressing: one `Grep class/def` at the start, one `Read offset+limit` before each edit (or a careful re-grep after any insert that shifts lines), five `Edit`s. The Grep/overview may need to be refreshed mid-session once line numbers drift. + +The hidden cost of the built-in workflow is not in any single call — it's the compounding re-Reads and anchor recomputation across a session. A single Read of a 1500-line file is ~40 kB of context; doing it four extra times across a session is a ~160 kB invisible tax that never shows up in a one-call comparison. + +**Intermediate output survives.** The overview I pulled at the start of this evaluation, the reference list for `Collector`, the type hierarchy for `BaseCollector`, and the signature table for `_collect` all remain valid now that I'm writing the report — I never had to regenerate them. A workflow based on line numbers would have had to regenerate its intermediate output after each of the ~15 edits I applied and reverted during the experiments. + +**Verdict (§6):** session-level efficiency scales with mutation rate; the more edits you plan, the more decisively symbolic addressing wins, and the effect is invisible on any single-call benchmark. + +--- + +## §7. Capabilities with no built-in equivalent + +1. **Cross-file atomic refactorings with automatic import maintenance.** `move` cleaned up an unused import in the source file as a side effect of moving the last user of that import. Built-ins have no equivalent — you'd have to notice. **Value: rare but high — this is the kind of cleanup that bit-rots across a large refactor.** + +2. **Resolution of third-party symbols from a usage site.** `find_declaration` + `find_symbol(search_deps=True)` reach into `site-packages` and `.pyi` stubs for the exact class used at a given line of code. **Value: a few times per session, saves 3–4 calls each.** + +3. **Transitive type hierarchy including external supertypes.** `type_hierarchy` returns sub- and super-chains in one call and crosses module boundaries and stub files. No text-search sequence can reproduce this in O(1). **Value: once or twice per unfamiliar codebase; near zero once you know it.** + +4. **Containing-symbol metadata on reference queries.** `find_referencing_symbols` returns each reference with its enclosing function/class name, making the result a navigation map rather than a line list. **Value: every time you need to understand *how* a symbol is used, not just *where*.** + +5. **Enforced safe-delete with usage-list refusal.** `safe_delete` refuses to remove a symbol that still has usages and returns the usage list. Built-ins cannot refuse — `Edit` applies whatever you send it. **Value: rare but prevents the worst-class mistake.** + +**Verdict (§7):** the capability deltas are real but concentrated in lower-frequency tasks; the one that shows up across *every* editing session is (1.5) — symbolic addressing as a property of the *editing* tools themselves, which I'm treating as the Tier-A efficiency win in §1 rather than a separate capability here. + +--- + +## §8. Where built-ins remain the right default + +- **Small anchored edits (≤ ~10 lines).** Task 7a: Edit's payload is ~10× smaller than `replace_symbol_body` for a one-line change because it sends only the two anchors, not the whole body. Frequency: extremely high — typo fixes, constant changes, error-message tweaks, single-line bug fixes. **Probably 30–50% of daily edits.** + +- **Free-text search for strings, log messages, magic constants, URLs.** Task 20: `Grep` is the only tool that can find a bare string across the repo. Serena's symbolic search expects an identifier, not a phrase. **Frequency: multiple times per session.** + +- **Non-code files** — changelogs, READMEs, YAML configs, notebooks. Task 19: `Read` is the tool. **Frequency: occasional but universal.** + +- **Doc and docstring sweeps after a code-level rename.** Serena's `rename_in_comments=True` catches Sphinx cross-references (verified in Task 10), but if your documentation lives outside of Python docstrings — `.md`, `.rst`, `.ipynb` — a text-based `Grep` sweep is the complementary step, not a Serena failure. **Frequency: every cross-file rename that touches a public API.** + +- **Single-file renames of distinctive identifiers.** Task 9: `Edit replace_all=true` is effectively tied with semantic rename; either works. **Frequency: common.** + +- **Quick one-shot explorations where you don't plan to edit the target.** If the workflow is "look at one function, answer a question, move on," `Read offset/limit` and `Grep` are as fast as semantic tools and don't require any address to be stable beyond the current call. + +These cases are not rare — collectively they probably cover more than half of the calls in a typical session, which is exactly why §1's weighting puts the quiet "addressing stability" win above the spectacular "cross-file refactoring" wins: the former touches every session, the latter touches a handful per week. + +**Verdict (§8):** built-ins are the right default for small edits, free-text search, non-code files, and docstring sweeps — roughly half of daily editing work. + +--- + +## §9. Usage rule for a developer with both toolsets + +Per-task decision rule, in priority order: + +1. **Small edit (≤ ~10 lines), known text anchor** → `Edit`. Payload is ~10× smaller than symbolic body replacement at this size. +2. **Medium or larger body rewrite, or structural insert** → `replace_symbol_body` / `insert_after_symbol`. 2–3× payload cut plus stable addressing for the next edit. +3. **Cross-file rename, move, or delete of a symbol** → `rename` / `move` / `safe_delete`, then a complementary `Grep` sweep of `.md`/`.rst`/`.ipynb` for any text-only references. The semantic tool covers code and Python docstrings in one atomic step; the Grep sweep handles external docs. +4. **Find callers of a class/function** → `find_referencing_symbols`. +5. **Find any mention of a name across the repo including docs** → `Grep`. +6. **Navigate into third-party library source** → `find_declaration` with a regex anchored at the use site, or `find_symbol(search_deps=True)`. +7. **Understand a type hierarchy** → `type_hierarchy`. +8. **Understand an unfamiliar file's structure** → `get_symbols_overview(depth=1)` for classes + methods, then `find_symbol(include_body=True)` for the methods you actually want to read. +9. **Chained edits to one file** → address every edit by `name_path`; do not mix in line-number or large-text-anchor edits unless you plan to re-Read before each one. +10. **Free-text search, non-code files, small single-line tweaks** → built-ins. + +**Verdict (§9):** Edit for small changes, symbolic body replacement for larger ones, semantic multi-file refactoring tools with a Grep doc-sweep on top, Grep/Read for text and non-code; the one habit that matters most is defaulting to name-path addressing whenever you expect more than one edit to the same file. diff --git a/docs/04-evaluation/results/020_codex_on_jbplugin-serena-evaluation.md b/docs/04-evaluation/results/020_codex_on_jbplugin-serena-evaluation.md new file mode 100644 index 00000000..2b091003 --- /dev/null +++ b/docs/04-evaluation/results/020_codex_on_jbplugin-serena-evaluation.md @@ -0,0 +1,136 @@ +# Evaluation by Gpt 5.4 in Codex CLI + +## Final Recommendation + +The recommendation is created with the [follow-up summary prompt](../011_followup-summary-prompt), everything else was created in one shot from the +[evaluation prompt](../010_evaluation-prompt). + +> "As a coding AI agent, I would ask my owner to add Serena because it lets me work with code as stable symbols instead of fragile text, turning the +refactors, lookups, and multi-file edits that usually feel risky and repetitive into precise IDE-backed operations I can trust." + +## 1. Headline + +Serena's largest added value is stable semantic addressing and IDE-backed refactoring. In this Java plugin repo, that meant a method/class could be targeted as `Logger/warning[0]`, `Symbol/safeDelete`, or `ProjectUtil` without line numbers, and cross-file rename/move/inline/delete operations were delegated to IntelliJ's model. + +High value, common: symbol lookup, method-body retrieval, stable name paths, and reference search. Frequency: many times per coding session. Value per hit: usually saves 1-3 reads/searches and avoids loading whole files. + +High value, less frequent but large: cross-file rename, move, safe delete, inline. Frequency: a few times per feature/refactor. Value per hit: saves roughly 5-20 calls and reduces missed import/call-site risk. + +Medium value: type hierarchy and external dependency declaration lookup. Frequency: occasional. Value per hit: turns "search and infer" into 1 semantic query; text search cannot truly reproduce transitive hierarchy or dependency source lookup without IDE/index/cache work. + +Low/no added value: config/docs/free-text search and tiny line edits. Frequency: common, but built-ins are already optimal. Value per hit for Serena: none or negative for small local edits. + +**Verdict:** Serena adds the most value whenever the task is about named code entities rather than text spans; the weighted daily win is stable symbol navigation, while the biggest per-hit win is IDE refactoring. + +## 2. Added Value By Area + +- Stable symbol navigation: showed on `Symbol.java` and `UIControlUtil.java`. Frequency: every non-trivial coding session. Value: saves 1-3 calls per lookup and avoids full-file reads. +- Symbol-scoped edits: `replace_symbol_body`, `insert_after_symbol`, and overload-specific targeting worked on `Logger/warning[0]`, `Logger/warning[1]`, and `Logger/logToToolWindow`. Frequency: several times per session. Value: small edits lose to text Edit, medium edits break even, whole-body edits save about 2x input payload. +- Cross-file refactors: renaming `Logger` to `SerenaLogger` updated 10 files plus the Java file rename; moving `ProjectUtil` into `service.endpoint` updated the package and removed the now-local import. Frequency: occasional. Value: saves roughly 10-20 manual reads/edits/verifications. +- Semantic relationships: `ToolWindowContent` references, `TypeHierarchy` subtypes, `SubtypeHierarchy` supertypes, and `Gson.fromJson` declaration came back as code entities. Frequency: occasional. Value: 1 call versus several searches plus inference. +- Built-in text work remains essential: `build.gradle.kts`, `/findSymbol`, `127.0.0.1`, and `FORM_INIT_DELAY_MILLIS` were naturally handled by `Read`/`rg`. Frequency: large share of daily work. Value: Serena adds nothing there. + +**Verdict:** Serena's contribution is not blanket speed; it removes repeated code-entity bookkeeping from ordinary navigation and almost all bookkeeping from real refactors. + +## 3. Detailed Evidence + +### 3.1 Code Understanding + +Semantic overview of `Symbol.java` returned a class tree with fields, methods, inner classes, and overload indexes in 1 call. Text equivalent was `rg "class |...\\(" Symbol.java`, which returned 100+ signature/comment hits and needed filtering. Follow-up semantic read of `UIControlUtil/findButton` was 1 call returning only the 10-line body; text follow-up needed locating the line then reading a slice. + +Payloads: semantic overview sent path/depth only and returned about 900 tokens; text grep returned about 2,000+ tokens for `Symbol.java`. Semantic method read returned about 90 tokens; text slice returned similar body tokens but required a locator step. + +**Verdict:** Use Serena for source structure and specific method bodies; use text only when the question is literally textual. + +### 3.2 References, Hierarchy, Dependencies + +`find_referencing_symbols` on `ToolWindowContent` returned 4 code uses: two subclass declarations and two parameters. `rg "ToolWindowContent"` returned 5 lines including the definition. For "who uses this in code," Serena was higher precision; for "where is this string mentioned," `rg` was the right tool. + +`type_hierarchy` on `TypeHierarchy` returned `SubtypeHierarchy` and `SupertypeHierarchy`; supertypes for `SubtypeHierarchy` returned `TypeHierarchy` and external `Object`. Text search found name matches plus false positives like `TypeHierarchyRequest`. + +`find_declaration` on `gson.fromJson(requestBody, requestClass)` resolved external `Gson/fromJson[0]` and returned the source body. Built-ins needed finding the Gradle dependency, locating `~/.gradle/.../gson-2.10.1-sources.jar`, listing/extracting `Gson.java`, then searching inside it. + +**Verdict:** Serena adds real semantic reach for "code relationships"; text search can find mentions, but not reliably answer relationship questions in one step. + +### 3.3 Edit Size Economics + +Small edit: rewriting `Logger.logToToolWindow` as an early return. Text patch sent a small old/new anchor, about 9 changed lines. Serena required sending the whole 11-line method body. Text was cheaper. + +Medium edit: rewriting most of `UIControlUtil.tryHandleDialogs`. Text patch sent about 14 old lines plus 20 new lines. Serena sent the full method body, including its comment, about 26 lines. Roughly even. + +Large edit: replacing `Symbol.safeDelete` implementation shape. With content-anchored Edit, a whole-body replacement would send old body plus new body, about 2x the new method payload. `replace_symbol_body` sent only the new body plus `Symbol/safeDelete`. + +**Verdict:** Use text Edit for 1-3 line changes, either tool for 10-30 line method rewrites, and Serena for whole-method replacements. + +### 3.4 Refactors + +Private rename: `Logger/logToToolWindow` to `appendToToolWindow` changed 7 occurrences in 1 semantic call plus optional `rg` verification. Manual path was `rg`, patch each occurrence, then `rg` verify: 3 calls and more payload. + +Multi-file rename: `Logger` to `SerenaLogger` changed 10 files and renamed the Java file. Manual path would be `rg`, read 10 files, rename/move file, edit imports/type names/constructors, verify with `rg`, and likely compile: roughly 14-20 calls. + +Move: moving `ProjectUtil` to `service.endpoint` moved the file, changed its package, and removed the import from `RefreshFileHandler` in 1 call. Manual path would coordinate filesystem move, package line, import removal/additions, and verification. + +Safe delete: `DebugUtil` had no references; safe delete returned `affected_references: []` and deleted it. Manual path is search, delete, verify. Inline: `ProjectUtil/getAbsolutePath` inlined into both call sites in 1 call. + +**Verdict:** Every multi-file or semantic refactor tested was a clear Serena win in call count, payload, and correctness surface. + +### 3.5 Session Effects + +I chained three edits in `Logger.java`: rename helper, replace `Logger/warning[0]`, insert after `Logger/error[0]`. The name paths survived earlier edits; no line recalculation was needed. A built-in line/slice workflow would need refreshed context after insertions because line numbers and nearby anchors shift. + +**Verdict:** Serena's stable identifiers compound across a session; the more edits you make in one file, the wider the gap gets. + +## 4. Token Efficiency + +The crossover is size-dependent. Small edit: text wins because it sends only a tiny anchor. Medium rewrite: near parity. Whole-body rewrite: Serena wins because it sends new body only, while content-anchored Edit sends old body plus new body. + +Forced reads matter: `find_symbol(...include_body=true)` avoided reading 573 lines of `UIControlUtil.java` and 1,295 lines of `Symbol.java`. Stable outputs also have longer shelf life: `Logger/warning[0]` remains useful after unrelated edits; line 52 does not. + +**Verdict:** Token economics favor built-ins for tiny local text edits and Serena for symbol retrieval, chained work, and whole-symbol replacement. + +## 5. Reliability + +Semantic matching distinguished overloads: `Logger/warning[0]` was `warning(String,Object...)`; `Logger/warning[1]` was `warning(String,Throwable)`. `rg "warning\\("` returned both and left disambiguation to the caller. + +Atomicity matters: semantic rename/move run as IDE refactorings, so a failure is not a half-finished sequence of 10 manual edits. Both toolsets still report mechanical success only; semantic intent still needs diff/build review. + +Transitive queries are where text cannot compete directly: hierarchy and external declaration lookup depend on IDE indexes and dependency sources, not just strings. + +**Verdict:** Correctness weight favors Serena for code-entity scope and refactoring atomicity, while text remains correct for text questions. + +## 6. Workflow Effects + +Over a multi-step session, Serena's intermediate artifacts stay usable: symbol trees, name paths, reference lists, and hierarchy nodes survive edits outside those symbols. Text outputs are often ephemeral: line numbers and byte offsets decay immediately after insertions/deletions. + +The practical effect is not just fewer calls; it is fewer re-reads. In the chained `Logger.java` session, Serena needed three edit calls. A text workflow would normally require initial reads/searches, edits, and refreshed context before later insertions. + +**Verdict:** Serena's session-level multiplier comes from not having to rediscover where code moved after each edit. + +## 7. No Built-In Equivalent + +- IDE semantic rename/move/inline/safe-delete: high value when refactoring, moderate frequency. Built-ins can approximate with many edits but cannot provide IDE refactoring semantics or atomicity. +- Type hierarchy including external `Object`: medium value, occasional. Text can search `extends`, but cannot transitively resolve hierarchy with dependency/stub awareness in one call. +- External dependency declaration resolution: medium value, occasional. Built-ins require build-file discovery and local source/cache spelunking. +- Overload/name-path targeting: high value in typed code, frequent enough to matter. Text can match names but cannot address `warning[0]` as a distinct method without manual signature reasoning. + +**Verdict:** Serena's unique capabilities are concentrated in IDE-indexed code intelligence and refactoring operations, not general file manipulation. + +## 8. Built-Ins As Default + +Use built-ins for non-code files, config, docs, and free text. `build.gradle.kts` was best read directly. `rg` was clearly right for `/findSymbol`, `127.0.0.1`, and `FORM_INIT_DELAY_MILLIS`. + +Use built-ins for tiny edits where a short unique anchor is obvious. Also always keep a post-refactor text sweep for comments, markdown, notebooks, generated files, and product strings; that is complementary verification, not a Serena weakness. + +Estimated share: built-ins remain best for maybe 40-60% of daily interactions because much coding work is still file/text/config/search. Serena dominates the code-symbol subset. + +**Verdict:** Start with built-ins for text and config; switch to Serena as soon as the noun in your task is a symbol. + +## 9. Practical Rule + +Reach for Serena when the task says class, method, overload, implementation, reference, hierarchy, rename, move, inline, delete, or "insert after this method." Reach for `rg`/Read/Edit when the task says string, config, docs, log message, URL, generated text, or "change these two lines." + +For edits: text Edit for 1-3 line tweaks; Serena `replace_symbol_body` for full methods/classes; semantic refactor tools for any rename/move/delete/inline that crosses call sites or imports. + +I restored all tracked edits after the experiments. The tracked working tree was clean after evaluation; the only baseline untracked entries were `.claude/` and `serena-evaluation-prompt.md`. I did not run the Gradle test suite because the task was an evaluation, not a product change. + +**Verdict:** With both toolsets installed, use built-ins for text and Serena for code entities; that rule captures almost all of the measured value without overthinking each call.