Skip to content

fix(infra): Telegram 文本报告改用 HTML 渲染并与 QQ 官方排版对齐 - #246

Merged
SXP-Simon merged 1 commit into
SXP-Simon:mainfrom
shiitin:fix/tg-text-report
Sep 26, 2026
Merged

SXP-Simon merged 1 commit into
SXP-Simon:mainfrom
shiitin:fix/tg-text-report

Conversation

@shiitin

@shiitin shiitin commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

📝 变更说明 (Summary)

修复 Telegram 文本报告中 Markdown 标记原样显示(群里看到 分析报告)的问题,并让 Telegram 文本报告复用 QQ 官方的 Markdown 排版(身份降级为昵称),使两端观感一致。共涉及 5 个文件、+273/−21。

🎯 变更背景与 STAR 描述

问题背景 (Problem)

三个叠加的症状,导致 Telegram 上的文本报告既「没渲染」又「排版走样」:

  1. Markdown 不渲染:报告文本自带 Markdown,而 TelegramAdapter.send_text 调用 client.send_message 时未传 parse_mode,Bot API 按纯文本发送 → 群内原样显示 ** 等标记。
  2. 重复的段落标题:Telegram 没有原生合并转发,format_forward_nodes_to_text 会把平台按段落切分的节点扁平化成一条消息;而每个节点都被命名为「分析报告」,扁平化时又逐段插入 [节点名] → 群里每段前面都重复一次「[分析报告]」。
  3. 两端排版不同:ReportDispatcher._dispatch_text 只对 qq_official 走 QQ 官方 Markdown 模块,其它平台走旧纯文本模板 —— 同一份数据在 QQ 与 Telegram 上章节不一致(Telegram 侧连「活跃时间分布」都没有),用户引用还会直接输出裸 ID(如 [8362414977])。

另外自查时发现 send_text 是所有文本的公共出口(图片标题、进度提示、命令回复都走它),若无差别做 Markdown 转换会误伤普通文本(详见「补充说明」)。

版本溯源:v5.6.3 的 send_text 同样只传 chat_id/text,属既有问题、并非本次重构引入。

解决措施 (Action)

  1. TelegramMessageConverter 新增 escape_html / to_telegram_html / strip_markdown / looks_like_markdown_report:加粗、 代码 转 / ,# 标题 转加粗行、- 列表 转 •、> 引用 转
    ,并转义 & < >(HTML 模式必须转义,否则发送直接失败);中标题(##/###)后补一个空行,对齐 QQ 客户端的段落间距;行内代码段先摘成占位符、其余转换做完再还原,避免 x 里的标记被当成加粗。
  2. TelegramAdapter.send_text 改用 parse_mode=「HTML」 发送,HTML 失败时自动降级为 strip_markdown 纯文本,保证不丢消息。同时加了硬门控:仅当文本确实含 Markdown 结构(** 或 # 标题行)时才走 HTML 模式,其余文本按原样纯文本发送。
  3. format_forward_nodes_to_text:同名节点不再逐段插入标题,正文自带 Markdown 标题时也不再补插「📊 分析报告」横幅;名称确实不同(典型的多来源转发)时仍保留逐段标题。
  4. 新增 QQOfficialMarkdownReportGenerator.generate_plain_markdown_report:复用 QQ 官方 Markdown 排版,身份由 <@user_id> 降级为昵称,且只替换明确的 ID 引用 —— [id]、<@id> 以及长度 ≥ 6 的纯数字裸 ID(平台用户 ID 的长度下限),不做昵称子串替换,也不会把「消息总数 123」这类统计值当成身份引用。
  5. ReportDispatcher 通过 _SHARED_MARKDOWN_PLATFORMS(当前为 {「telegram」})决定是否走该排版;由于该方法不在 IReportGenerator 契约内,调用前做能力检查,实现缺失时回退 generate_text_report,自定义生成器不会抛 AttributeError。

验证效果 (Result)

  • 真机 AstrBot 4.28.1 + 真实 Telegram 群触发文本报告,Bot API 返回的 Message.entities 含 24 个 bold 与 4 个 blockquote(消息 id 23;另有自动识别的 1 个 URL 与 1 个 EMAIL),群内加粗、• 列表、引用块均正常渲染。
  • 报告正文中的 a<b & c 正确转义为字面量,未触发 HTML 解析错误。
  • 章节与 QQ 官方一致(含活跃时间分布),用户引用显示为昵称而非裸 ID,逐段「[分析报告]」前缀消失。
  • 回归门禁:pytest tests/ 382 passed;ruff check / ruff format --check 无告警;pyright 与未打补丁的 v5.6.5 基线逐条比对为 0 新增诊断。

🔗 关联 Issue (Related Issue)

  • Fixes #(暂无对应 issue,如有请补充)

📦 变更类型 (Type of Change)

  • ✨ feat: 新增功能 / 新增 WebUI 模块
  • 🐛 fix: 修复 Bug / 异常处理
  • 📚 docs: 文档或注释更新
  • 💄 style / ♻️ refactor: 报告排版与渲染路径统一
  • ⚡ perf: 性能优化
  • 🧪 test: 增加或修改单元测试 / 冒烟测试
  • 🔧 chore / 👷 ci: 构建流程 / 工作流 / 依赖包版本变动

📋 提交前自检清单 (Checklist)

  • 提交规范:Commit Message 符合 Conventional Commits & STAR 标准(scope = infra,正文含中文「问题/解决措施/效果」三点论)
  • 跨端契约一致性:本 PR 不涉及后端 API 变更
  • 前端验证:本 PR 不涉及 WebUI(未改动 dashboard/ 与 pages/)
  • 后端质量:ruff check / ruff format --check 无告警,pytest tests/ 382 passed
  • 本地门禁自检:本机未安装 lefthook,已逐项运行等价命令(ruff check、ruff format --check、npx pyright、node scripts/verify-commit.js),结果全部通过

📌 补充说明

  • 回归自查(重要):send_text 是所有文本(含图片标题、进度提示)的公共出口,因此逐条核对过「非报告文本」不被误改。实测中抓到两种真实误伤并已消除:

    • 234=24 → 234=24
    • C:/logs/.txt 与 .py → 被当作一对斜体跨段吞掉

    处理方式:彻底移除斜体转换(报告本身只用 ** 加粗,收益低误伤高),并加上 looks_like_markdown_report 硬门控。图片标题、进度提示、数学式子、通配符路径、- / > 开头的普通文本现均保持原样纯文本发送,行为与改动前逐字一致。

  • payload 兼容性:QQ 生成器多读了 topic.contributor_ids,该字段在 SummaryTopic 中有 default_factory=list、反序列化走 .get(「contributor_ids」, []) → 老 checkpoint 不会 AttributeError。

  • QQ 官方路径行为不变:mention_style 默认为 qq,仅 Telegram 走昵称降级分支。

  • 与其它 PR 的关系:本 PR 与「群列表恒空」修复都改 telegram_adapter.py,但落点分别在约 118 行与约 380 行,互不重叠,任意顺序合并均可(已验证两两合并与三合一合并均无冲突)。

  • 顺带发现(本 PR 未修改,建议单独修):qq_official_markdown.render_identity_text 在把昵称替换为 @提及 时使用裸子串替换 source.replace(name, replacement) —— 群内存在单字昵称(如「夜」)时会把「深夜/熬夜/夜间」全部插上 @;同一昵称对应多个 ID 时替换值为空字符串,会把昵称整个删掉。该函数对用户 ID 的替换有边界保护,唯独昵称这一行没有。Telegram 不经过这段代码,故未动。


🔍 评审回应(Sourcery 首轮)

四条意见逐条处理,代码均已改动并复测:

  1. generate_shared_markdown_text_report 不在 IReportGenerator 契约内,自定义实现会抛 AttributeError —— 属实,已按建议改为 getattr 能力检查 + 回退 generate_text_report。
    补充一点供参考:紧邻的 QQ 官方分支调用的 generate_qq_official_markdown_report 同样不在该契约内,属仓库既有写法,本 PR 未动它(避免扩大改动面)。
  2. 身份替换范围比声称的宽(ID 为 123 时「消息总数 123」会被替换成昵称) —— 属实,已收窄为只替换 [id]、<@id> 与长度 ≥ 6 的纯数字裸 ID,并加了常量 _MIN_BARE_ID_LENGTH 与注释说明(QQ 号 6~11 位、Telegram ≤10 位,6 位下限即为此设)。
  3. 加粗转换早于代码转换, x 会被当成嵌套加粗 —— 属实,已改为先把代码段摘成占位符、其余转换做完再还原;实测 x 现在输出 x。
  4. 类注释里仍写着 斜体,但实现已无斜体转换 —— 属实(斜体正是本 PR 主动移除的,用于避免 234、*.py 被误渲染),注释已删掉斜体说法并注明移除原因。

复测证据:容器内对四个点逐项定点验证通过(代码段保护、短 ID 不替换、长 ID / [id] / <@id> 正常替换、注册表接口缺失打 ERROR 并回退),真机真实群文本报告 entities 仍为 24 bold + 4 blockquote,仓库 382 项单测全过。

Summary by Sourcery

统一 Telegram 文本报告的富文本渲染与 QQ 官方排版,并增强发送失败降级和兼容性处理。

Bug Fixes:

  • 修复 Telegram 文本报告中的 Markdown 标记不渲染、重复段落标题、章节缺失和用户 ID 裸露等问题。
  • 为 Telegram HTML 渲染失败提供安全的纯文本降级,避免网络类错误导致重复发送。

Enhancements:

  • 统一 Telegram 与 QQ 官方的 Markdown 报告排版,并将用户身份在 Telegram 中降级显示为昵称。
  • 改进 Telegram 长文本按行切分及转发节点排版,避免破坏格式并消除同名报告标题重复。
  • 增强身份引用替换边界和报告生成器兼容性,减少统计数字误替换及自定义实现缺失接口导致的异常。

Chores:

  • 改进 Telegram 群列表 KV 回退的接口校验与错误日志。

@sourcery-ai

sourcery-ai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

本 PR 将 Telegram 文本报告改为按需转换的 HTML 富文本,增加 HTML 失败后的纯文本兜底,消除扁平化报告的重复标题,并复用 QQ 官方 Markdown 报告排版;无提及能力时通过严格的 ID 引用识别显示用户昵称,同时兼容缺少可选接口的自定义报告生成器。

Sequence diagram for Telegram Markdown report delivery

sequenceDiagram
    participant Dispatcher as ReportDispatcher
    participant Generator as ReportGenerator
    participant Adapter as TelegramAdapter
    participant Telegram as TelegramBotAPI

    Dispatcher->>Generator: generate_shared_markdown_text_report(analysis_result)
    Generator-->>Dispatcher: Markdown report
    Dispatcher->>Adapter: send_text(group_id, text)
    Adapter->>Adapter: looks_like_markdown_report(text)
    alt Markdown report detected
        Adapter->>Adapter: to_telegram_html(text)
        Adapter->>Telegram: send_message(text, parse_mode=HTML)
        alt HTML delivery fails
            Telegram-->>Adapter: error
            Adapter->>Adapter: strip_markdown(text)
            Adapter->>Telegram: send_message(text)
        end
    else Ordinary text
        Adapter->>Telegram: send_message(text)
    end
Loading

File-Level Changes

Change Details Files
将 Telegram 报告 Markdown 转换为安全的 Telegram HTML,并保留普通文本行为及发送兜底。
  • 新增 HTML 转义、标题/列表/引用/加粗/代码转换与 Markdown 检测逻辑。
  • 仅对检测为报告 Markdown 的文本使用 HTML 模式,失败时降级为去标记纯文本;普通文本原样发送。
  • 保护行内代码内容并移除高误伤的斜体转换。
src/infrastructure/platform/adapters/telegram_adapter.py
src/infrastructure/platform/adapters/telegram_message_converter.py
修正 Telegram 扁平化转发报告的标题重复问题。
  • 同名节点合并正文并取消逐段标题。
  • 多来源节点继续保留段落标题;正文已有 Markdown 标题时不再添加额外报告横幅。
src/infrastructure/platform/adapters/telegram_message_converter.py
让 Telegram 文本报告复用 QQ 官方 Markdown 报告结构,并将用户身份降级为昵称。
  • 新增共享 Markdown 报告生成入口,覆盖 QQ 官方报告的章节、主题、称号、金句等排版。
  • 仅替换明确的用户 ID 引用及长度达标的数字裸 ID,避免误改普通统计数字或昵称子串。
  • 通过平台集合选择共享排版,并对缺少可选生成方法的自定义生成器回退至普通文本报告。
src/infrastructure/reporting/dispatcher.py
src/infrastructure/reporting/generators.py
src/infrastructure/reporting/qq_official_markdown.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@shiitin shiitin closed this Sep 26, 2026
@shiitin shiitin reopened this Sep 26, 2026
@shiitin
shiitin marked this pull request as ready for review September 26, 2026 13:43

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 4 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="src/infrastructure/reporting/dispatcher.py" line_range="554-560" />
<code_context>
             )
+        elif self._uses_shared_markdown_report(platform_id):
+            # 无 @ 提及能力的平台复用 QQ 官方的 Markdown 排版(身份降级为昵称)
+            text_report = self.report_generator.generate_shared_markdown_text_report(
+                analysis_result
+            )
         else:
</code_context>
<issue_to_address>
**issue (bug_risk):** The dispatcher calls `generate_shared_markdown_text_report` for Telegram, but that method is not part of the `IReportGenerator` contract, so a configured custom or test implementation of the interface raises `AttributeError` instead of dispatching the report.

**Triggers:** When the dispatcher is constructed with an `IReportGenerator` implementation other than the concrete `ReportGenerator` class.

**Suggested fix:** Add the method to `IReportGenerator`, or use a capability check and fall back to `generate_text_report` when the implementation does not provide it.

```suggestion
        elif self._uses_shared_markdown_report(platform_id):
            # 无 @ 提及能力的平台复用 QQ 官方的 Markdown 排版(身份降级为昵称)
            shared_markdown_generator = getattr(
                self.report_generator, "generate_shared_markdown_text_report", None
            )
            if callable(shared_markdown_generator):
                text_report = shared_markdown_generator(analysis_result)
            else:
                text_report = self.report_generator.generate_text_report(analysis_result)
        else:
            text_report = self.report_generator.generate_text_report(analysis_result)
```
</issue_to_address>

### Comment 2
<location path="src/infrastructure/reporting/qq_official_markdown.py" line_range="198-202" />
<code_context>
+            name = names[user_id]
+            source = source.replace(f"[{user_id}]", name)
+            source = source.replace(f"<@{user_id}>", name)
+            source = re.sub(
+                rf"(?<![A-Za-z0-9_[<]){re.escape(user_id)}(?![A-Za-z0-9_>])",
+                name,
+                source,
+            )
</code_context>
<issue_to_address>
**issue (bug_risk):** The supposedly explicit-reference-only replacement also replaces any standalone occurrence of a known user ID, including an ordinary statistic or other report value. For example, if a user ID is `123`, a message count or topic text containing standalone `123` is replaced with that user's nickname.

**Triggers:** When a report value happens to equal a known user ID as a standalone token.

**Suggested fix:** Restrict replacement to explicit `[id]` and `<@id>` forms, or introduce a report-specific reference format instead of treating every standalone ID as an identity reference.

```suggestion

```
</issue_to_address>

### Comment 3
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="88-89" />
<code_context>
+        escaped = cls._MD_SUBHEADING_RE.sub(r"<b>\1</b>\n", escaped)
+        escaped = cls._MD_QUOTE_RE.sub(r"<blockquote>\1</blockquote>", escaped)
+        escaped = cls._MD_BULLET_RE.sub("• ", escaped)
+        escaped = cls._MD_BOLD_RE.sub(r"<b>\1</b>", escaped)
+        return cls._MD_CODE_RE.sub(r"<code>\1</code>", escaped)
+
+    @classmethod
</code_context>
<issue_to_address>
**nitpick (bug_risk):** Bold conversion runs before code conversion, so Markdown markers inside inline code are interpreted as HTML formatting. A literal code span such as `` `**x**` `` becomes nested `<b>` markup inside `<code>` instead of displaying the literal `**x**` content.

**Triggers:** When report content contains `**...**` inside an inline-code span.

**Suggested fix:** Protect code spans before applying bold, heading, quote, or list transformations, then restore them after the other Markdown conversions.
</issue_to_address>

### Comment 4
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="31" />
<code_context>
 class TelegramMessageConverter:
     """Telegram 消息转换与昵称自愈器。"""

+    # Telegram 文本渲染:报告文本带 markdown 语法(**加粗** / *斜体* / `代码`),
+    # 而 Bot API 默认按纯文本发送,导致群内看到的是 ** 原样字符。
+    # 这里把 markdown 转成 Telegram HTML 模式可识别的标签,并转义 HTML 保留字符。
+    _MD_BOLD_RE = re.compile(r"\*\*(.+?)\*\*", re.S)
+    _MD_CODE_RE = re.compile(r"`([^`\n]+)`")
</code_context>
<issue_to_address>
**nitpick:** The class comment states that the converter handles `*斜体*`, but the implementation no longer has an italic conversion rule, so the comment documents behavior that the code does not provide.

**Suggested fix:** Remove the italic claim from the comment or document the intentional decision that italic conversion was removed.

```suggestion
    # Telegram 文本渲染:报告文本带 markdown 语法(**加粗** / `代码`),
```
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 2 findings to address first, and this changes the automatically generated reports sent to Telegram users, including identity rendering and HTML conversion. If the conversion or name mapping is wrong, incorrect content can be sent externally and remains in those messages after a revert; reverting only fixes future reports.

Blocking findings: src/infrastructure/reporting/dispatcher.py:560, src/infrastructure/reporting/qq_official_markdown.py:202


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread src/infrastructure/reporting/dispatcher.py
Comment thread src/infrastructure/reporting/qq_official_markdown.py Outdated
Comment thread src/infrastructure/platform/adapters/telegram_message_converter.py Outdated
Comment thread src/infrastructure/platform/adapters/telegram_message_converter.py Outdated
@shiitin
shiitin marked this pull request as draft September 26, 2026 13:47
@shiitin
shiitin force-pushed the fix/tg-text-report branch 2 times, most recently from 32ed8b9 to 9c568f1 Compare September 26, 2026 14:08
@shiitin
shiitin marked this pull request as ready for review September 26, 2026 14:12
sourcery-ai[bot]
sourcery-ai Bot previously approved these changes Sep 26, 2026

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Approved.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

@shiitin

shiitin commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@sourcery-ai review

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 1 issue

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="107-108" />
<code_context>
+    @classmethod
+    def strip_markdown(cls, text: str) -> str:
+        """去除 markdown 标记,得到纯文本(HTML 渲染失败时的兜底)。"""
+        plain = cls._MD_BOLD_RE.sub(r"\1", text)
+        return cls._MD_CODE_RE.sub(r"\1", plain)
+
     @staticmethod
</code_context>
<issue_to_address>
**issue (broader_impact):** 纯文本降级先执行加粗正则、再执行代码正则,因此代码段中的 `**` 会先被当成报告加粗标记处理;例如 `` `**x**` `` 会降级为 `x`,丢失代码内容中的字面量星号,而 HTML 路径已经专门通过占位符避免了这一问题。

**Triggers:** 当报告包含行内代码中的 `**`,且 HTML 发送失败时。

**Suggested fix:** 与 `to_telegram_html` 一样先摘取代码段,再处理加粗和其他 Markdown 标记,最后恢复代码内容。
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 1 finding to address first, and incorrect Markdown-to-HTML conversion or identity replacement could send malformed or misleading report content to Telegram users, and an HTML-send exception could potentially result in duplicate delivery if the first request succeeded. Reverting stops future messages but does not retract reports already sent outside the team.

Blocking findings: src/infrastructure/platform/adapters/telegram_message_converter.py:108


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread src/infrastructure/platform/adapters/telegram_message_converter.py Outdated
@shiitin

shiitin commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@sourcery-ai review

@sourcery-ai
sourcery-ai Bot dismissed their stale review September 26, 2026 14:32

Sourcery withdrew this approval because the latest commits introduced blocking findings.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 3 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="src/infrastructure/platform/adapters/telegram_adapter.py" line_range="404-413" />
<code_context>
+                    message_thread_id=thread_id,
+                    reply_to_message_id=reply_id,
+                )
+            except Exception as html_error:
+                logger.warning(
+                    f"[Telegram] HTML 模式发送失败,降级为纯文本: {html_error}"
+                )
+                await client.send_message(
+                    chat_id=chat_id,
</code_context>
<issue_to_address>
**issue (bug_risk):** When Telegram accepts the HTML request but the client receives a timeout or other transport exception, the code sends the same report again as plain text, producing duplicate report messages.

**Triggers:** When the HTML send succeeds remotely but its response is lost or times out.

**Suggested fix:** Retry only on errors that prove the request was not accepted, or use an idempotency mechanism instead of unconditionally sending the fallback.
</issue_to_address>

### Comment 2
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="74-102" />
<code_context>
+    def to_telegram_html(cls, text: str) -> str:
</code_context>
<issue_to_address>
**issue (bug_risk):** Reports longer than the adapter's 4000-character chunks are converted one chunk at a time, so Markdown delimiters split across chunks are never paired; the affected chunks retain raw `**` or backticks instead of rendering the formatting.

**Triggers:** When a long report is split inside a bold or inline-code span by `send_forward_msg`.

**Suggested fix:** Split on safe Markdown boundaries before conversion, or preserve formatting state across chunks and convert the complete report before performing a size-aware split.
</issue_to_address>

### Comment 3
<location path="src/infrastructure/reporting/qq_official_markdown.py" line_range="199-214" />
<code_context>
+        统计数字当成身份引用(ID 为 123 的用户不会让"消息总数 123"变成昵称)。
+        """
+        source = str(text or "")
+        for user_id in sorted(names, key=len, reverse=True):
+            name = names[user_id]
+            source = source.replace(f"[{user_id}]", name)
+            source = source.replace(f"<@{user_id}>", name)
+
+        bare_ids = [
+            user_id
+            for user_id in names
+            if user_id.isdigit() and len(user_id) >= cls._MIN_BARE_ID_LENGTH
+        ]
+        for user_id in sorted(bare_ids, key=len, reverse=True):
+            source = re.sub(
+                rf"(?<![A-Za-z0-9_[<]){re.escape(user_id)}(?![A-Za-z0-9_>])",
+                names[user_id],
+                source,
+            )
+        return source.strip()
+
</code_context>
<issue_to_address>
**issue (bug_risk):** Identity replacement is performed sequentially on the already-modified source, so a nickname inserted for one user is scanned again for every other user's bare numeric ID and can be partially replaced or corrupted.

**Triggers:** When a collected nickname contains another known numeric user ID, especially in Chinese text where the regex treats digits adjacent to non-ASCII characters as standalone IDs.

**Suggested fix:** Protect explicit identity replacements with placeholders, or perform all bare-ID substitutions against the original text in one regex pass and restore the protected names afterward.
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 3 findings to address first, and the new Markdown-to-HTML path and identity substitution can send malformed or incorrect report content, including wrong nickname replacements or formatting, to Telegram users. Reverting prevents future messages but cannot retract reports already delivered outside the team; the impact is bounded to those sent messages.

Blocking findings: src/infrastructure/platform/adapters/telegram_adapter.py:413, src/infrastructure/platform/adapters/telegram_message_converter.py:102, src/infrastructure/reporting/qq_official_markdown.py:214


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread src/infrastructure/platform/adapters/telegram_adapter.py
Comment thread src/infrastructure/platform/adapters/telegram_message_converter.py
Comment thread src/infrastructure/reporting/qq_official_markdown.py Outdated
问题:报告文本自带 markdown(含转发节点 **[分析报告]** 段落),而 TelegramAdapter.send_text 调用
client.send_message 时未传 parse_mode,Bot API 按纯文本发送,群内看到 ** 等标记原样字符;同时非
QQ 官方平台走另一套纯文本模板,用户引用直接输出裸 ID(如 [8362414977]),章节也比 QQ 官方少了
「活跃时间分布」,两端体验完全不一致。

解决措施:
- TelegramMessageConverter 新增 escape_html / to_telegram_html / strip_markdown / looks_like_markdown_report:
  **加粗**、`代码` 转 <b>/<code>,# 标题转加粗行、- 列表转 •、> 引用转 <blockquote>,并转义 & < > 三个字符;
  行内代码段在 HTML 与纯文本两条路径里都先摘成占位符、其余转换做完再还原,避免 `**x**` 这类代码内容被误处理;
- TelegramAdapter.send_text 改用 parse_mode="HTML" 发送,HTML 发送失败自动降级为去标记纯文本,确保不丢消息;
  但仅当文本确实含 Markdown 结构(** 加粗或 # 标题行)时才走 HTML 模式 —— 图片标题、进度提示等普通文本
  仍按原样纯文本发送,且不再做斜体转换,避免把正文里的 2*3*4、*.py 这类字符误渲染成标签;
- 新增 QQOfficialMarkdownReportGenerator.generate_plain_markdown_report:复用 QQ 官方 Markdown 排版,
  身份由 <@user_id> 降级为昵称展示;只替换 [id] / <@id> 以及长度 >= 6 的纯数字裸 ID,既不误伤「深夜」
  这类含短昵称的正文,也不会把「消息总数 123」这种统计值当成用户 ID;
- ReportDispatcher 对不具备 @ 提及能力的平台(当前为 Telegram)改走该排版;该方法不在 IReportGenerator
  契约内,调用前做能力检查,实现缺失时回退 generate_text_report,自定义生成器不会再抛 AttributeError;
- send_text 的纯文本降级只在「HTML 被 Telegram 拒绝」(BadRequest / can't parse entities)时触发;
  网络/超时类错误可能已投递成功,重发会造成群内重复报告,改为直接返回失败交给上层;
- 长报告分段(>4000 字符)改为按行边界切分,避免把跨行的 **加粗** / `代码` 标记从中间切断;
- 转发节点扁平化不再逐段插入 **[分析报告]** 标题:平台按段落切分时所有节点同名,逐段加标题
  会让群里每段前面都重复一次「[分析报告]」,现在同名节点直接拼正文,正文自带 Markdown 标题时
  也不再多插一次「📊 分析报告」横幅(多来源转发等名称不同的场景仍保留逐段标题)。

效果:真机 AstrBot 4.28.1 + 真实 Telegram 群,发送后 Bot API 返回的 entities 含 4 个 blockquote 与
24 个 bold(另有自动识别的 URL / EMAIL),群内加粗、列表与引用块正常渲染,报告里的 a<b & c 正确转义
为字面量;文本报告章节与 QQ 官方一致(含活跃时间分布),用户引用显示为昵称而非裸 ID,且不再出现逐段
「[分析报告]」前缀;代码段保护、短 ID 不替换、生成器接口缺失回退三项回归点已在容器内定点复测通过
(HTML 与纯文本降级两条路径都保留代码段字面量);身份替换改为对原文单次遍历,替换出的昵称不会被二次扫描;
长文本切分保持内容无损(逐段拼接等于原文);仓库 382 项单测全过。
@SXP-Simon
SXP-Simon merged commit 49d768a into SXP-Simon:main Sep 26, 2026
4 checks passed
SXP-Simon added a commit that referenced this pull request Sep 26, 2026
问题:缺少对 PR #245(Telegram 注册表群组发现与 TargetResolver 注入)、PR #246(Telegram Markdown 转 HTML 与转发排版)、PR #247(QQ 官方适配器未初始化容错)以及统一 Markdown 报告链路的独立单元测试。

解决措施:新增 test_telegram_message_converter.py 与 test_telegram_group_discovery.py,补充 QQ 官方未初始化单测,并同步更新多平台 Markdown 报告测试用例。

效果:单元测试用例从 383 个扩充至 393 个,全量执行 100% 通过,为跨平台与报告重构提供完整覆盖防护。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants