<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>AI 落地与企业应用雷达</title><link>https://radar.yuedu.biz</link><description>本期有 7 条生产或规模化信号达到 A/B 级。高价值信息的共同点，是同时披露业务场景、运行机制与可复核结果。</description><item><title>LinkedIn 让客服 Agent 在生产中持续自我演进</title><link>https://arxiv.org/html/2608.10224</link><guid>https://arxiv.org/html/2608.10224</guid><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate><description>LinkedIn 在真实企业客服流量上部署可持续演进的 Agent，把提示词、检索与评估做成可版本化闭环。两周随机 A/B 测试显示，QA 自助率提升 9.0 个百分点、取消流程自助率提升 4.8 个百分点、路由准确率提升 30.6 个百分点，并包含分阶段发布、回归检查与回滚机制。</description></item><item><title>美妆品牌用 GEA 将 80 个 SKU 制作周期从 16 周压缩至 3 周</title><link>https://www.tezign.com/industries/beauty-brand-gea-sku-production</link><guid>https://www.tezign.com/industries/beauty-brand-gea-sku-production</guid><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate><description>某美妆品牌将 GEA 企业智能体投入 80 个 SKU 的详情页生产工作流，制作周期从预计 16 周缩短至 3 周，A/B 测试变体由平均 1 个增加到 4 个；案例来自供应商官方页面，客户匿名，结果仍需独立核验。</description></item><item><title>PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage</title><link>https://arxiv.org/abs/2608.24040</link><guid>https://arxiv.org/abs/2608.24040</guid><pubDate>Tue, 25 Aug 2026 03:57:56 GMT</pubDate><description>Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its…</description></item><item><title>Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation</title><link>https://arxiv.org/abs/2608.23992</link><guid>https://arxiv.org/abs/2608.23992</guid><pubDate>Tue, 25 Aug 2026 02:33:44 GMT</pubDate><description>Large language model (LLM) agents invoke external tools to retrieve and reason over information beyond pretrained knowledge. The Model Context Protocol (MCP) standardizes how such tools are surfaced, and a proxy MCP ser…</description></item><item><title>快手 AgentX 将推荐实验变成自迭代生产线</title><link>https://arxiv.org/pdf/2606.26859</link><guid>https://arxiv.org/pdf/2606.26859</guid><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><description>快手把 AI AgentX 生产部署到主站推荐和本地生活业务三周，三个 Agent worker 将 374 个想法收敛成 10 次可上线实验，报告 8 倍并发、相对人工工程师 3.7 倍业务价值、0.561% 用户时长提升，以及超过人民币 1 亿元年化收入。论文没有提供可比基线。</description></item><item><title>PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?</title><link>https://arxiv.org/abs/2608.26882</link><guid>https://arxiv.org/abs/2608.26882</guid><pubDate>Thu, 27 Aug 2026 09:40:11 GMT</pubDate><description>Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: c…</description></item><item><title>Constraint-Guided Enterprise Data Mapping with Large Language Models</title><link>https://arxiv.org/abs/2608.24218</link><guid>https://arxiv.org/abs/2608.24218</guid><pubDate>Tue, 25 Aug 2026 08:23:45 GMT</pubDate><description>Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or granularity mismatches. Manual matching is still common in practice, but does not scale as schemas and providers evolve…</description></item><item><title>Salesforce 与 Anthropic 发布 Claudeforce</title><link>https://www.salesforce.com/ap/news/press-releases/2026/08/27/salesforce-and-anthropic-announce-claudeforce-the-1-ai-meets-the-1-ai-crm</link><guid>https://www.salesforce.com/ap/news/press-releases/2026/08/27/salesforce-and-anthropic-announce-claudeforce-the-1-ai-meets-the-1-ai-crm</guid><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><description>Salesforce 与 Anthropic 将企业 AI Claude 接入 Salesforce 数据、业务规则和工作流，并以 37 项销售技能开始试点。Salesforce 同时披露内部生产部署的 Slackbot 带来 810 万小时年化生产力收益；外部客户效果尚未披露，因此属于产品发布与内部采用证据。</description></item><item><title>腾讯把推荐系统 Agent 的自主权限制在决策点</title><link>https://arxiv.org/html/2608.11241</link><guid>https://arxiv.org/html/2608.11241</guid><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate><description>系统在腾讯三个推荐业务线生产部署 78 天，记录 1,624 次工具调用。两个业务线的新业务接入时间由约 14 天缩短至约 3 天；论文明确说明该数字来自作者回忆、没有受控前序基线，应作为方向性案例而非普遍结论。</description></item><item><title>京东用 LLM/VLM 建成工业级商品知识平台</title><link>https://doi.org/10.48550/arxiv.2606.28070</link><guid>https://doi.org/10.48550/arxiv.2606.28070</guid><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><description>京东在电商核心业务的数百亿 SKU 上生产部署 LLM/VLM 商品知识平台，每天处理数亿次商品更新。论文披露知识生产精确率 94.2%、召回率 82.8%，商品信息质量问题下降 37%，核心属性自动填充率超过 80%，商品创意优化点击率提升约 9%。</description></item><item><title>Learning Generalizable Behaviors for Terminal Agents</title><link>https://arxiv.org/abs/2608.22631</link><guid>https://arxiv.org/abs/2608.22631</guid><pubDate>Wed, 26 Aug 2026 21:18:13 GMT</pubDate><description>Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users&#x27; daily workflows. Reinforcement learning (RL) is a key technique for improving their capabi…</description></item><item><title>Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav</title><link>https://arxiv.org/abs/2608.24764</link><guid>https://arxiv.org/abs/2608.24764</guid><pubDate>Tue, 25 Aug 2026 16:03:02 GMT</pubDate><description>Large language model agents are moving beyond conventional retrieval-augmented generation toward direct interaction with external corpora. Direct Corpus Interaction (DCI) keeps the full corpus accessible, yet reachable…</description></item><item><title>Benchmarking the Titans: A Multi-Dimensional Empirical Evaluation of LLM Code Generation Quality in the .NET Ecosystem</title><link>https://arxiv.org/abs/2608.22529</link><guid>https://arxiv.org/abs/2608.22529</guid><pubDate>Sun, 23 Aug 2026 17:59:19 GMT</pubDate><description>Evaluating Large Language Model (LLM) code generation quality requires examining not just whether the generated code is correct, but whether it is maintainable, efficient, and stylistically sound, all of which are quali…</description></item><item><title>Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency</title><link>https://arxiv.org/abs/2608.26885</link><guid>https://arxiv.org/abs/2608.26885</guid><pubDate>Thu, 27 Aug 2026 09:42:39 GMT</pubDate><description>Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abst…</description></item><item><title>Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems</title><link>https://arxiv.org/abs/2608.25920</link><guid>https://arxiv.org/abs/2608.25920</guid><pubDate>Wed, 26 Aug 2026 15:33:47 GMT</pubDate><description>As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing…</description></item><item><title>When &quot;Must&quot; Becomes &quot;Maybe&quot;: Constraint Weakening in LLM Agent Workflows</title><link>https://arxiv.org/abs/2608.24569</link><guid>https://arxiv.org/abs/2608.24569</guid><pubDate>Tue, 25 Aug 2026 13:51:52 GMT</pubDate><description>Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets…</description></item><item><title>LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks</title><link>https://arxiv.org/abs/2608.23200</link><guid>https://arxiv.org/abs/2608.23200</guid><pubDate>Tue, 25 Aug 2026 12:35:20 GMT</pubDate><description>Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet succe…</description></item><item><title>AI-Assisted Extraction of Follow-up Observations from GCN Circulars in Astro-COLIBRI</title><link>https://arxiv.org/abs/2608.23270</link><guid>https://arxiv.org/abs/2608.23270</guid><pubDate>Mon, 24 Aug 2026 13:57:30 GMT</pubDate><description>We present a new Astro-COLIBRI component that converts free-text GCN Circulars into structured, event-linked follow-up records and combines them with structured reports submitted directly by the community. A continuousl…</description></item><item><title>The Reverse Big Push: Generative AI and Self-Fulfilling Automation</title><link>https://arxiv.org/abs/2608.25602</link><guid>https://arxiv.org/abs/2608.25602</guid><pubDate>Wed, 26 Aug 2026 10:23:02 GMT</pubDate><description>Generative AI relocates the fixed cost of automation. A model provider pays to train a frontier system, while a downstream firm rents capability by usage; the same firm must carry a continuing payroll to supply a human-…</description></item><item><title>Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays</title><link>https://arxiv.org/abs/2608.22068</link><guid>https://arxiv.org/abs/2608.22068</guid><pubDate>Sat, 22 Aug 2026 18:18:07 GMT</pubDate><description>Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations, provide opportunities to enhance energy production capacity and increase fault tolerance. The development…</description></item><item><title>Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models</title><link>https://arxiv.org/abs/2608.27165</link><guid>https://arxiv.org/abs/2608.27165</guid><pubDate>Thu, 27 Aug 2026 14:17:14 GMT</pubDate><description>Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows. Existing output-stage uncertainty metrics can…</description></item><item><title>Towards a universal meta-optics solver via large language models</title><link>https://arxiv.org/abs/2608.26417</link><guid>https://arxiv.org/abs/2608.26417</guid><pubDate>Wed, 26 Aug 2026 21:35:14 GMT</pubDate><description>Metasurface design increasingly requires fast models that can operate across structurally distinct device families, rather than retraining a separate surrogate for every geometry class. Conventional neural network surro…</description></item><item><title>Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence</title><link>https://arxiv.org/abs/2608.25869</link><guid>https://arxiv.org/abs/2608.25869</guid><pubDate>Wed, 26 Aug 2026 14:41:24 GMT</pubDate><description>Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, whe…</description></item><item><title>LMSM: LLM Security Framework Inspired by Linux Security Modules</title><link>https://arxiv.org/abs/2608.25697</link><guid>https://arxiv.org/abs/2608.25697</guid><pubDate>Wed, 26 Aug 2026 12:13:05 GMT</pubDate><description>Large language models (LLMs) are increasingly deployed with layered defenses, yet malicious prompts can still bypass them. Interpretability methods can expose model-internal signals along the generation path that could…</description></item><item><title>TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving</title><link>https://arxiv.org/abs/2608.25523</link><guid>https://arxiv.org/abs/2608.25523</guid><pubDate>Wed, 26 Aug 2026 08:33:10 GMT</pubDate><description>Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memor…</description></item><item><title>A Comparative Evaluation of Digitization Pipelines for Historiographical Sources</title><link>https://arxiv.org/abs/2608.24976</link><guid>https://arxiv.org/abs/2608.24976</guid><pubDate>Tue, 25 Aug 2026 16:03:01 GMT</pubDate><description>Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora p…</description></item><item><title>Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model</title><link>https://arxiv.org/abs/2608.24215</link><guid>https://arxiv.org/abs/2608.24215</guid><pubDate>Tue, 25 Aug 2026 08:21:58 GMT</pubDate><description>Large language model (LLM)-based multi-agent social simulation has demonstrated compelling results, but Agentopia was evaluated with 100 agents over 10 simulated years using Qwen3.5-397B-A17B, leaving the behavior of re…</description></item><item><title>TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers</title><link>https://arxiv.org/abs/2608.23763</link><guid>https://arxiv.org/abs/2608.23763</guid><pubDate>Mon, 24 Aug 2026 18:54:42 GMT</pubDate><description>The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised…</description></item><item><title>Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling</title><link>https://arxiv.org/abs/2608.23653</link><guid>https://arxiv.org/abs/2608.23653</guid><pubDate>Mon, 24 Aug 2026 11:50:07 GMT</pubDate><description>AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents different requirements from general-purpose code generation in software engineering, because correctness depends not o…</description></item><item><title>LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications</title><link>https://arxiv.org/abs/2608.23058</link><guid>https://arxiv.org/abs/2608.23058</guid><pubDate>Mon, 24 Aug 2026 10:02:39 GMT</pubDate><description>Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting a…</description></item></channel></rss>
