前沿专利观察 / Patent Frontier Watch

第005期:AI前沿:大模型治理、RAG与模型验证

本期重写为逐件研读版,重点不是罗列AI热词,而是看每件专利如何把模型、数据、验证、安全、算力或场景任务写成可执行的技术链条。

为什么选这些专利

本期选择标准不是随机编号,也不是标题看起来热门,而是每件专利都对应一个可被拆解的工程问题:它必须能说明一个具体痛点,并且在公开文本中给出结构、流程、材料组合、数据处理或控制策略。下面按逐件研读方式展开。

1. Detection of hallucinations in large language model responses

代表专利:US-2025225337-A1。申请人/权利人:GOOGLE LLC (US)。优先权日:2024/01/09;公开日:2025/07/10。

它真正发现的问题

这件专利的问题落在“Detection of hallucinations in large language model responses”这个具体AI场景:Implementations described herein relate to detecting hallucinations in responses generated by large language models (LLMs). 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:Implementations described herein relate to detecting hallucinations in responses generated by large language models (LLMs). A natural language (NL) based input associated with a client device may be received. A first LLM response may be generated based on processing the NL based input using an LLM. Based on processing the first LLM response, it may be determined whether the first LLM response contains at least one hallucination. Responsive to determining that the first LLM response contains at least one hallucinat...

给我们的启示

启示是,大模型治理专利要写清检测对象、触发条件、拦截动作和审计记录。企业布局时应保护模型外层的安全流程,而不只是模型本身。

2. Personalized retrieval-augmented generation system

代表专利:US-12373506-B1。申请人/权利人:DROPBOX INC (US)。优先权日:2024/01/23;公开日:2025/07/29;授权日:2025/07/29。

它真正发现的问题

这件专利的问题落在“Personalized retrieval-augmented generation system”这个具体AI场景:The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating personal responses through retrieval-augmented generation. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating personal responses through retrieval-augmented generation. In particular, the disclosed systems can generate a query embedding from a query generated by an entity and determine data context specific to the entity by comparing the query embedding with a plurality of vectorized segments of content items associated with the entity. The disclosed systems can provide the data context to a large language model a...

给我们的启示

启示是,RAG和企业AI的壁垒常在数据链路:检索什么、如何验证、怎样回填业务系统、错误如何处理,都是可保护点。

3. Ai hallucination and jailbreaking prevention framework

代表专利:US-2025045531-A1。申请人/权利人:Unum Group (US)。优先权日:2023/08/02;公开日:2025/02/06。

它真正发现的问题

这件专利的问题落在“Ai hallucination and jailbreaking prevention framework”这个具体AI场景:The disclosed embodiments include systems and methods configured to provide a Generative AI framework that uses the power of multiple LLMs by separating the generative aspect into multiple distinct large language models. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:The disclosed embodiments include systems and methods configured to provide a Generative AI framework that uses the power of multiple LLMs by separating the generative aspect into multiple distinct large language models. In some disclosed embodiments, a first large language model evaluates an input prompt and transforms it if needed (e.g., in a first processing stage of the framework); a second large language model performs a generative function based on an input prompt it receives from the first large language mo...

给我们的启示

启示是,大模型治理专利要写清检测对象、触发条件、拦截动作和审计记录。企业布局时应保护模型外层的安全流程,而不只是模型本身。

4. Machine learning model verification for assessment pipeline deployment

代表专利:US-2022129787-A1。申请人/权利人:PAYPAL INC (US)。优先权日:2020/10/27;公开日:2022/04/28。

它真正发现的问题

这件专利的问题落在“Machine learning model verification for assessment pipeline deployment”这个具体AI场景:There are provided systems and methods for machine learning model verification for assessment pipeline deployment. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:There are provided systems and methods for machine learning model verification for assessment pipeline deployment. A service provider may provide AI hosting platforms that allow for clients, customers, and other end users to upload AI models for execution, such as machine learning models. A user may utilize one or more user interfaces to provide model data and files, such as model artifacts, model requirements, and model test data. Thereafter, a model deployer may validate that the AI hosting platform has the requ...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

5. Method for certification of periodically adapted machine learning models

代表专利:EP-4379602-A1。申请人/权利人:FREQUENTIS AG (AT)。优先权日:2022/11/29;公开日:2024/06/05。

它真正发现的问题

这件专利的问题落在“Method for certification of periodically adapted machine learning models”这个具体AI场景:[Translated] A computer-implemented method (10) for verifying machine learning models in safety-critical environments is disclosed. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:[Translated] A computer-implemented method (10) for verifying machine learning models in safety-critical environments is disclosed. The computer-implemented method (10) comprises determining (S120) technical characteristics (25) and operational characteristics (30) for a selected operational machine learning model (15) in a simulation environment (20), continuously testing (S130) the selected operational machine learning model (15), generating (S150) a number n of optimized machine learning models (15a-15n), selec...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

6. Large language model regulation systems and methods

代表专利:US-2025045596-A1。申请人/权利人:INTUIT INC (US)。优先权日:2023/07/31;公开日:2025/02/06。

它真正发现的问题

这件专利的问题落在“Large language model regulation systems and methods”这个具体AI场景:At least one processor may receive a query response generated by a query machine learning (ML) model, wherein the query response is generated in response to a query from a client device. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:At least one processor may receive a query response generated by a query machine learning (ML) model, wherein the query response is generated in response to a query from a client device. The at least one processor may generate an evaluated likelihood of the query response being found in a training data set comprising known valid data, wherein the generating is performed using an evaluation ML model. The at least one processor may determine that the evaluated likelihood indicates the query response is likely to inc...

给我们的启示

启示是,大模型治理专利要写清检测对象、触发条件、拦截动作和审计记录。企业布局时应保护模型外层的安全流程,而不只是模型本身。

7. Object Classification in Image Data Using Machine Learning Models

代表专利:US-2018150713-A1。申请人/权利人:SAP SE (DE)。优先权日:2016/11/29;公开日:2018/05/31。

它真正发现的问题

这件专利的问题落在“Object Classification in Image Data Using Machine Learning Models”这个具体AI场景:Combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each p... 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:Combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object. The image data within each bounding polygon that is determined to encapsulate an object can then be provided to...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

8. Using a large language model for code generation for network analytics with coding hints

代表专利:US-2025060949-A1。申请人/权利人:CISCO TECH INC (US)。优先权日:2023/08/16;公开日:2025/02/20。

它真正发现的问题

这件专利的问题落在“Using a large language model for code generation for network analytics with coding hints”这个具体AI场景:In one implementation, a device pauses generation of computer code by a language model. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:In one implementation, a device pauses generation of computer code by a language model. The device matches a block of the computer code to a hint regarding a portion of the block of computer code. The device inserts the hint into the computer code. The device resumes generation of the computer code by the language model, wherein the language model uses the hint to generate a new portion of the computer code.

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

9. Utilizing machine learning models to synthesize perturbation data to generate perturbation heatmap graphical user interfaces

代表专利:US-12374429-B1。申请人/权利人:RECURSION PHARMACEUTICALS INC (US)。优先权日:2023/09/14;公开日:2025/07/29;授权日:2025/07/29。

它真正发现的问题

这件专利的问题落在“Utilizing machine learning models to synthesize perturbation data to generate perturbation heatmap graphical user interfaces”这个具体AI场景:The present disclosure relates to systems, non-transitory computer-readable media, and methods for embedding perturbation data via a machine learning model and filtering, aligning, and aggregating the embeddings to generate a genome-wide perturbation database for real-time generation of perturbation heatmaps. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:The present disclosure relates to systems, non-transitory computer-readable media, and methods for embedding perturbation data via a machine learning model and filtering, aligning, and aggregating the embeddings to generate a genome-wide perturbation database for real-time generation of perturbation heatmaps. In particular, in one or more embodiments, the disclosed systems can receive a plurality of perturbation images portraying cells from a plurality of wells corresponding to a plurality of cell perturbations. F...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

10. Ideographic contrastive autoencoder for large language model fine-tuning

代表专利:US-12572783-B1。申请人/权利人:STRAVA INC (US)。优先权日:2024/10/31;公开日:2026/03/10;授权日:2026/03/10。

它真正发现的问题

这件专利的问题落在“Ideographic contrastive autoencoder for large language model fine-tuning”这个具体AI场景:Ideographic contrastive autoencoder for large language model fine-tuning is disclosed, including: obtaining a set of user activities according to a specified task; obtaining respective sets of input features from the set of user activities; using an encoder network of an autoencoder to encode the respective sets of input features into a set of words; prompt. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:Ideographic contrastive autoencoder for large language model fine-tuning is disclosed, including: obtaining a set of user activities according to a specified task; obtaining respective sets of input features from the set of user activities; using an encoder network of an autoencoder to encode the respective sets of input features into a set of words; prompting a machine learning model to perform the specified task using the set of words, wherein the machine learning model has been fine-tuned using a custom lexicog...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

11. Personalized communication in a digital therapy platform

代表专利:US-12397198-B1。申请人/权利人:SWORD HEALTH S A (PT)。优先权日:2024/02/23;公开日:2025/08/26;授权日:2025/08/26。

它真正发现的问题

这件专利的问题落在“Personalized communication in a digital therapy platform”这个具体AI场景:An example digital therapy platform is disclosed that provides personalized, real-time feedback to patients during sessions. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:An example digital therapy platform is disclosed that provides personalized, real-time feedback to patients during sessions. The platform collects real-time performance data during the sessions, including metrics such as range of motion, pelvic floor movements, exercise completion rates, and the accuracy of movements. The digital therapy platform may also retrieve historical data from past sessions to provide a comprehensive overview of the patient's context. The digital therapy platform dynamically generates stru...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

12. Watermark unit for a data processing accelerator

代表专利:US-11645586-B2。申请人/权利人:BAIDU USA LLC (US) KUNLUNXIN TECH BEIJING COMPANY LIMITED (CN)。优先权日:2019/10/10;公开日:2023/05/09;授权日:2023/05/09。

它真正发现的问题

这件专利的问题落在“Watermark unit for a data processing accelerator”这个具体AI场景:In one embodiment, a computer-implemented method performed by a data processing (DP) accelerator, the method includes receiving, at the DP accelerator, first data representing a set of training data from a host processor and performing training of an artificial intelligence (AI) model based on the set of training data within the DP accelerator. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:In one embodiment, a computer-implemented method performed by a data processing (DP) accelerator, the method includes receiving, at the DP accelerator, first data representing a set of training data from a host processor and performing training of an artificial intelligence (AI) model based on the set of training data within the DP accelerator. The method further includes implanting, by the DP accelerator, a watermark within the trained AI model and transmitting second data representing the trained AI model having...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

13. Enhanced grading and feedback assistant system for handwritten student work

代表专利:US-2025239171-A1。申请人/权利人:GOODNOTES LTD (GB)。优先权日:2024/01/23;公开日:2025/07/24。

它真正发现的问题

这件专利的问题落在“Enhanced grading and feedback assistant system for handwritten student work”这个具体AI场景:This disclosure describes systems, methods, and devices for artificial intelligence-based grading and feedback of digitally entered handwritten characters into a device. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:This disclosure describes systems, methods, and devices for artificial intelligence-based grading and feedback of digitally entered handwritten characters into a device. A method may include converting, using a first device, a computer-readable document with questions into a digital worksheet comprising teacher layers and student layers; detecting, using a first machine learning model trained to categorize questions and generate answer zones for the questions, answer zones; generating first updated student layers ...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

14. Systems and methods for engineering protein activity

代表专利:US-2024013854-A1。申请人/权利人:AETHER BIOMACHINES INC (US)。优先权日:2022/01/10;公开日:2024/01/11。

它真正发现的问题

这件专利的问题落在“Systems and methods for engineering protein activity”这个具体AI场景:The present disclosure provides systems and methods for engineering protein activity. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:The present disclosure provides systems and methods for engineering protein activity. In an aspect, described herein are new predictive models that predict protein activity from its amino acid sequence (e.g., using a trained machine learning model and a physics-based simulation model). In another aspect, new prescriptive models identify one or more candidate proteins (e.g., for use in the predictive model) that can use a multi-objective search and optimization algorithm. In another aspect, predictive and prescript...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

15. Robust neural network learning system

代表专利:US-12243295-B2。申请人/权利人:GM GLOBAL TECH OPERATIONS LLC (US)。优先权日:2022/03/31;公开日:2025/03/04;授权日:2025/03/04。

它真正发现的问题

这件专利的问题落在“Robust neural network learning system”这个具体AI场景:A system comprises a computer including a processor and a memory. The memory includes instructions such that the processor is programmed to: receive intermediate concept constraints at a neural network and train the neural network with training data, training labels, and the at least one of the data constraint, the feature constraint, or the intermediate co... 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:A system comprises a computer including a processor and a memory. The memory includes instructions such that the processor is programmed to: receive intermediate concept constraints at a neural network and train the neural network with training data, training labels, and the at least one of the data constraint, the feature constraint, or the intermediate concept constraint.

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

16. System and method for player reidentification in broadcast video

代表专利:US-12288342-B2。申请人/权利人:STATS LLC (US)。优先权日:2019/02/28;公开日:2025/04/29;授权日:2025/04/29。

它真正发现的问题

这件专利的问题落在“System and method for player reidentification in broadcast video”这个具体AI场景:A system and method of re-identifying players in a broadcast video feed are provided herein. 研读重点是它如何把模型能力约束到可验证的数据处理、判断或部署流程中。

它的技术方案

公开摘要显示,方案的具体构成是:A system and method of re-identifying players in a broadcast video feed are provided herein. A computing system retrieves a broadcast video feed for a sporting event. The broadcast video feed includes a plurality of video frames. The computing system generates a plurality of tracks based on the plurality of video frames. Each track includes a plurality of image patches associated with at least one player. Each image patch of the plurality of image patches is a subset of the corresponding frame of the plurality of ...

给我们的启示

启示是,AI专利应把模型放回具体任务闭环,写清数据、模型、验证、反馈和部署载体之间的关系。

本期结论

这些AI专利共同说明,真正可布局的不是“用了大模型/机器学习”,而是模型外层的验证、安全、数据入口、硬件执行和具体业务闭环。

资料来源 / Sources

本期依据 PubChem patent JSON、Google Patents 公开页面和原公开链接研读整理。正式FTO或授权稳定性判断仍需进一步核对官方登记簿、审查历史、同族和权利要求全文。