OIML BULLETIN - 2026 - VOLUME LXVII - NUMBER 3
f o c u s p a p e r
Preserving Legal-Metrology Integrity in AI-Assisted Logging
Auditable Linking of Recognition and Weighing
Kamacho Scale Co., Ltd. , Japan
Citation: T. Kamada 2026 OIML Bulletin LXVII(3) 20260305
Abstract
Industrial records must answer three questions – what, how much, and when. Traditionally, “how much” has been the domain of legally verified weighing instruments, while “what” and “when” have depended on human inspection and manual entry. The rise of AI image recognition promises automation of “what”, while weighing instruments continue to deliver mass indications verified under the applicable national legal-metrology regime. Yet two challenges have hindered deployment of fully automated records suitable for subsequent audit and, where accepted under national law, legally significant use. First, training an AI to relate an image to a mass indication has required humans to assemble each (image, mass)-pair by hand – a foundational bottleneck of industrial AI. Second, and – more consequential from a legal-metrology standpoint – temporal proximity of the image and the mass indication cannot, by itself, prove that the two refer to the same physical object; the evidence of how they were linked has been missing from industrial records. This article proposes a framework that bridges recognition and weighing by generating, preserving, and rendering auditable the linking evidence of every event. The output is a formally defined record of image, mass indication, time, and linking evidence, hash-chained for tamper detection – suitable both as a tamper-evident, auditable log with an explicit audit trail from identification to measurement, and as paired training data for AI carrying a scale-derived mass label. The mass value itself remains that of a weighing instrument verified under the applicable national legal-metrology regime on the basis of OIML R 76: image and AI processing complement, but do not intrude upon, the metrologically traceable mass indication. The framework is illustrated by a pilot deployment on a recycling sorting line.
1. Introduction
Weighing has stood at the foundation of trade and industry for millennia. Its authority in commerce, regulation, and science derives from a chain of metrological trust – one that the International Organization of Legal Metrology (OIML) has, over decades, codified and internationalized through Recommendations such as R 76 for non-automatic weighing instruments [1]. When properly verified, the scale does not merely produce a number: it produces a number backed by an auditable trail linking that measurement to national standards and, ultimately, to the SI. Throughout this article, “verified” refers to verification under the applicable national legal-metrology regime on the basis of the relevant OIML Recommendation.
Yet the reliable recording of industrial events requires more than mass alone. Any complete record must answer three questions: what passed through this point, how much of it, and when. In many legacy or mixed manual–digital workflows, these three attributes are met by heterogeneous means: mass by the verified scale, identity by human inspection, and timestamps by manual data entry. The framework proposed here targets, in particular, objects that cannot carry persistent barcode or RFID identifiers. Only the mass indication is subject to formal legal-metrology control. The identity determination is subject to human judgement, and the timestamp to human error, omission, or tampering.
The maturation of artificial intelligence, particularly in image recognition [2], has opened the possibility of automating the identity determination. This has led, over the past decade, to numerous proposals for fusion systems combining image and weight to produce complete, automated records. In practice, however, deployment has been slow. This article argues that two challenges – one economic, one evidentiary – have slowed deployment, and that both can be addressed by a single framework that binds recognition and weighing through recorded linking evidence, while preserving the legal-metrology status of the underlying weighing instrument.
The article is structured as follows. Section 2 characterizes the current state of industrial recording. Section 3 introduces the two challenges and situates them in prior work. Section 4 describes the proposed framework, including a formal record definition (Figure 1) and its relation to legal-metrology governance. Section 5 presents a pilot deployment and two further application classes, and Section 6 discusses implications for the international weighing community. We round off with our Conclusions in Section 7.
2. Recording under Legal Metrology: The Current State
An industrial record is a factual assertion, made now, about a past event, intended to be trusted later. For this trust to be well-founded, each element of the assertion must be traceable to a reliable source. The current state is uneven.
What? Object identity – in most industrial applications the object’s category, rather than an individual identity – is typically determined by a human inspector or, increasingly, by an AI image classifier [2]. Preliminary internal measurements at the author’s company indicated that image-only mass estimation is substantially less reliable than domain-specific object classification, owing to variations in density, moisture, packing, and shape.
How much? Mass is measured by weighing instruments conforming to OIML Recommendations – most commonly R 76 for non-automatic instruments [1], R 51 for automatic catchweighing [5], or R 134 for weigh-in-motion of road vehicles [6] – with legal-metrology accuracy classes determined by verification-scale-interval requirements. Mass is thus obtained from an instrument subject to the applicable legal-metrology controls and maximum-permissible-error requirements; where an application-specific uncertainty evaluation has been performed, the associated uncertainty may be recorded separately in accordance with the Guide to the Expression of Uncertainty in Measurement (GUM) [3].
When? In such workflows, timestamps are commonly generated by human entry into a paper form or an information system. Unlike mass or identity, timestamps are seldom verified: they are subject to omission, delay, and deliberate manipulation, and rarely carry a chain of custody.
The composite record thus rests on a heterogeneous foundation: rigorously traceable in one attribute, technologically credible in a second, and administratively fragile in the third. This is the state that fusion of AI and weighing has long promised to improve.
3. Two Challenges: Manual Pairing and the Evidentiary Gap
Two challenges have kept fusion from advancing beyond limited pilots.
The manual pairing challenge (economic)
To train an AI model that relates images to mass (or that classifies objects into mass-relevant categories), a body of labelled (image, mass) pairs must be assembled. Historically, this has meant that a human photographs an object, places it on a scale, records the mass indication, labels the image, and saves the pair – repeated thousands or tens of thousands of times per site. A typical site therefore requires months of labour and considerable expense before any AI can begin to learn, and the exercise must be repeated for every new region, contaminant, species, or vehicle type. This is a foundational bottleneck of industrial AI deployment.
The evidentiary gap (metrological)
A less-discussed but more consequential challenge arises from the temporal structure of real deployments. In practice, image capture and mass measurement seldom coincide: a worker picks up an item from a conveyor, carries it to a bin, and drops it; a robotic arm suctions an object and places it on a weigh plate; a truck is imaged in motion and weighed only after coming to rest. Under these conditions, timestamp proximity alone cannot prove that a particular image and a particular mass indication refer to the same physical object. Yet, absent such proof, the fused record is not an auditable assertion but only a plausible association. From the standpoint of legal metrology this is a significant concern: a record whose composition cannot be verified after the fact is inadequate as evidence for trade, regulation, or dispute resolution.
Relation to prior work
The ingredients of a solution exist in separate literatures: multi-object tracking maintains object identity across frames [8]; provenance models such as W3C PROV formalize how a data item was derived [9]; hash-chained timestamping secures the integrity of event logs [10]; and chain-of-custody terminology and models are standardized in ISO 22095 [11], albeit for physical supply chains – the term is used here by analogy, the digital derivation and integrity aspects of the proposed record being more directly captured by data-provenance and audit-trail concepts. On the basis of a review of the legal-metrology, industrial-AI, and data-provenance literature available to the author as of mid-2026, however, the author is not aware of a published record format that retains the image–mass pairing itself as auditable evidence; in the fusion systems known to the author, the pairing is an internal computation whose evidence is discarded once the record is written. Both challenges must be addressed if fusion is to yield records that satisfy legal-metrology expectations; neither is addressed by improving image recognition or scale accuracy in isolation.
4. The Framework: Recorded Linking Evidence
4.1 Principle
We propose a framework whose organizing principle is that the connection between an image and a mass indication is not asserted by temporal proximity, but established by recording the evidence of how, in each event, the two were bound – and by making that evidence a first-class, auditable component of the record itself (Figure 1a). The framework is a generalizable design pattern connecting two activities: recognition (image-based identification by AI) and weighing (mass measurement by a verified instrument).
Figure 1. The framework. (a) Recognition and weighing are bridged by linking evidence recorded during the object’s transit from camera view to scale view. (b) Each event yields a hash-chained record in which the linking evidence is a first-class, auditable component.
4.2 Methods that Generate Linking Evidence
Linking evidence is produced by maintaining and logging the continuity of the object during its transit between camera view and scale view. Three classes of methods are distinguished by how the evidence is generated.
Physical methods rely on controlled physical possession or mechanically observable containment of the object during transit, producing a kinematic evidence log. Examples include a robot’s gripper or suction head, whose trajectory from grasp to release on the scale is logged; trays or baskets whose contents are auditable at boundaries; and singulated conveyor belts whose constant velocity yields a deterministic position–time mapping. A worker’s hand may also provide physical continuity; where the hand state is detected by vision, as in the pilot of Section 5, the resulting method is classified as hybrid. In that instance, continuity is detected by a four-state detector – idle, hand entered, object held, hand exited holding – whose final transition fires the pairing trigger; the stored image is taken from a frame buffer at
Logical methods rely on image processing to track the object across frames or across cameras, producing a vision-based evidence log. Examples include multi-object tracking on an open conveyor [8], logging bounding-box positions and identities across frames; re-identification across non-overlapping cameras, logging match scores between images; and trajectory prediction using conveyor-speed priors.
Hybrid methods combine physical and logical evidence for cross-checked provenance, in which the two evidence logs corroborate each other.
The framework is method-agnostic. What matters, from an evidentiary standpoint, is that each pairing of image and mass indication is accompanied by recorded, verifiable evidence of how the link was established – in each class, with a defined evidence payload. The choice of method follows the site, not the framework – a property that enables application across recycling lines, livestock corridors, logistics gates, and beyond.
4.3 The Record: Formal Definition
Each event i yields the record
Ri = (Ii, wi, ti, Ei, hi),
where Ii is a reference to the immutable image together with its digest di = SHA-256(image bytes) (identifying “what”); wi the mass indication obtained from a weighing instrument verified under the applicable national legal-metrology regime on the basis of the relevant OIML Recommendation, stored together with the instrument identifier, verification status, and applicable metrological characteristics (quantifying “how much”); ti a timestamp synchronized to a designated time source – NTP in the pilot deployment, a trusted time service in production (recording “when”); and Ei = (mi, Li, ci) the linking evidence, comprising the method identifier mi ∈ {physical, logical, hybrid}, the evidence log Li (a gripper trajectory, a state sequence, or tracking data), and a normalized pairing confidence ci ∈ [0, 1]. Pairings whose confidence falls below a site-specific threshold are rejected or referred to human review; for ci to support probabilistic claims it must be calibrated against observed mispairing rates as part of site validation – an uncalibrated score supports triage only. Moreover, confidence values are method-specific and not directly comparable across methods; an interoperable schema therefore also records method-specific quality metrics and a disposition status (accepted, rejected, or human-reviewed) for each pairing. The chain value hi = SHA-256(canon(Ii, wi, ti, Ei) ∥ hi−1), computed over a canonical serialization of the record content – fixed field order, number representation, and character encoding – binds each record to its predecessor and provides a tamper-evident integrity chain [10]: any subsequent modification breaks the chain and remains detectable, provided a trusted chain anchor is retained (Figure 1b). The linking evidence is thus not an internal computation hidden from the user: a verifier presented with a suspect record can examine Ei and – where the AI model, software versions, and configuration recorded with the event are available – re-run the linking procedure. The record supports subsequent examination of the association rather than constituting proof of it.
The five-tuple is a conceptual minimum. An interoperable record standard additionally carries a record identifier and sequence number; a schema version; algorithm identifiers for hashing and signature; separate UTC timestamps for image capture, stable indication, linking decision, and record commitment, each with a clock identifier and synchronization error bound; camera, instrument, and gateway identifiers; the unit and gross/net/tare designation with a stability flag; the instrument’s accuracy class, Max, Min, and verification scale interval e; a reference to the verification certificate and its validity; AI model, software, and configuration versions; a method-specific schema for the evidence log; content-addressed digests for the evidence log, AI model snapshot, configuration, and calibration data; and explicit representation of missing data, exceptions, and manual interventions.
Hash chaining detects tampering but does not by itself prevent it. Digital signatures over chain segments, managed keys, a trusted time source, and external anchoring are therefore necessary conditions for legally meaningful tamper evidence – part of the record standard itself rather than optional implementation detail – consistent with the software provisions of OIML D 31 [7]. Their detailed specification is beyond the scope of this article.
4.4 Relation to Legal-Metrology Governance
The framework operates downstream of the weighing indication, under explicit architectural constraints: communication from the instrument to the recording system is one-way, with no feedback path from the AI components to any weighing function; the received indication is stored unmodified, together with the displayed value, stability status, unit, and gross/net/tare designation; communication errors are recorded as exceptions; and each record carries the instrument identifier and a reference to the validity of its verification. AI-derived quantities (e.g., mass estimated from an image) are labelled as such and kept distinct from the metrologically traceable value. Where an application-specific uncertainty evaluation has been performed, its result is recorded separately in accordance with GUM [3].
R 76 does not specifically address AI-assisted evidentiary linking between recognition data and weighing results. OIML D 31:2023 [7], however, explicitly recognizes dynamic software modules whose parameters may evolve in use – including machine-learning processes – together with the notion of a snapshot of such a module’s state. Whether an external recording system of the kind proposed here is itself legally relevant will depend on the system configuration and on national law, particularly where the records are used for trade, statutory reporting, or enforcement. How such downstream multimodal records should be treated for legally or regulatorily significant purposes remains, in the author’s view, a governance question that merits attention in the OIML technical committees. Indeed, the OIML Digitalization Task Group is already preparing guidance on AI in legal metrology [13]; the record structure proposed here may offer a concrete use case for that ongoing work. Emerging guidance on AI risk management, such as the NIST AI Risk Management Framework [12], may usefully inform that discussion.
5. Applications
The framework has been piloted in the first of the following application classes; the second and third are at the design or early-deployment stage. Each class is characterized by a different choice of linking-evidence method.
Recycling sorting (piloted)
The pilot operates on a sorting line at the author’s company handling plastic, metal, paper, and e-waste. Each collection bin is placed on a verified weighing instrument, cameras observe the conveyor, and contaminants are removed both by workers and by a collaborative robot placed upstream of the manual station – an arrangement in which whatever the current model misses is picked by the worker, and each such pick automatically becomes a training sample for precisely the classes on which the model is weakest. Erroneous picks by the worker introduce label noise; because each sample carries its linking evidence and stored image, such samples can be audited and removed during curation. Linking evidence is hybrid, combining the worker’s physical possession of the object with vision-based detection of the hand-state sequence of Section 4.2 [4]. The mass assigned to an individual pick is derived from two stable indications acquired immediately before and after the deposit; both original indications, their stability flags, and the resulting difference are retained, and the legal-metrology status of the original indications is distinguished explicitly from that of the externally calculated difference. The applicable instrument classification and the result-acceptance procedure differ between the worker-operated and robot-operated modes and are determined under the applicable national regime. Category labels derive from the sorting action itself – the identity of the bin into which the object is deposited – rather than from manual annotation. From normal operation alone, the pilot produced an operational mass log and two site-specific models – a contaminant classifier and an image-based mass estimator – trained exclusively on framework records (Figure 2). The pilot demonstrated that framework records can be generated during routine operation and used to train site-specific models; no claim of validated model performance is made in this article, and the pilot is presented as an illustrative deployment rather than a controlled performance study. Detailed model performance, throughput, and labour-saving figures are commercially sensitive and are therefore outside the scope of this article. Independent quantitative validation remains an area for future work.
Figure 2. The pilot sorting line: (a) the site-specific contaminant detector, trained exclusively on records generated during normal operation, running on the conveyor (detections overlaid; white rectangle marks the detail region); (b) enlarged detail.
Livestock weight monitoring (early stage)
In pig-barn corridors, a scale embedded beneath a passage and an overhead camera together yield per-animal weight history without catching or handling the animals. Linking evidence is logical: cross-frame re-identification provides auditable evidence that the imaged animal and the weighed animal are the same. Over months, an AI trained on these records may estimate weight from image alone, extending weight-monitoring coverage beyond the physical footprint of the scale.
Overload screening in logistics (early stage)
Roadside cameras paired with verified weighing data produce per-vehicle load records. Linking evidence is logical and operates over road geometry: licence-plate identity and the vehicle’s trajectory between camera and weighing station serve as the verifiable trace. Over time, the AI may learn to estimate cargo mass from imagery alone, supporting pre-screening at locations without weighing equipment and directing selected vehicles to a verified weighing station. The image-derived estimate serves screening only: any legal action rests exclusively on a verified result obtained under the applicable national regime—a static weighbridge based on R 76 [1] or a weigh-in-motion instrument conforming to R 134 [6], as applicable – consistent with the distinction drawn in Section 4.4. Licence-plate imagery constitutes personal data in many jurisdictions; retention limits, access control, and data minimization are therefore part of the deployment design.
6. Implications for the International Weighing Community
Trade transparency and consumer protection
Records generated by the framework carry an explicit audit trail between identification and measurement. Disputes over what was weighed – historically settled by testimony or by expensive re-inspection – may instead be settled by inspection of the linking evidence embedded in the record itself.
Regulatory compliance
In sectors subject to statutory reporting (waste manifests, agricultural output declarations, transport compliance), the framework produces records whose fitness for regulatory submission is a property of construction, not of after-the-fact assembly.
The role of the weighing instrument in the AI era
The framework re-articulates the role of the weighing instrument. Traditionally, the instrument has been valued as the anchor of trust for a measured quantity. In the framework, it becomes additionally the anchor of trust for a labelled training sample destined for an industrial AI. The scale, in this sense, becomes not only the source of measurement but also the source of ground truth for the mass label; semantic category labels derive from an independent annotation source, such as the operator’s sorting action, the destination-bin identifier, or a separate labelling process.
Standardization opportunities
The value of the framework grows with interoperability. If manufacturers of weighing instruments across the OIML community adopt a common record format – including a common schema for linking evidence – records generated at one site can be aggregated across sites, and AI models trained on distributed evidence can be verified against a common evidentiary standard. Such standardization lies well within the OIML’s tradition of harmonizing metrological practice internationally, and the author invites the community to consider it.
7. Conclusion
This article has proposed a framework that bridges recognition and weighing not by temporal coincidence but by recording the evidence of how, in each event, an image was linked to a mass indication. The resulting record – image, mass indication, time, and linking evidence – serves simultaneously as a tamper-evident, auditable log with an explicit audit trail from identification to measurement, and as a paired multimodal training sample carrying a scale-derived mass label (semantic category labels require an independent annotation source). The framework is method-agnostic and applies wherever a scale and a camera observe the same physical object. The mass value itself remains that of a verified instrument, under the architectural constraints of Section 4.4.
Two contributions flow from this construction. The evidentiary contribution is that industrial records generated within the framework carry, by construction, auditable evidence of how the pairing between identification and measurement was established. The economic contribution is that the dedicated manual image–mass pairing step that has long bottlenecked industrial AI deployment can be avoided or substantially reduced by generating labelled records during normal operation. The principal limitation of the present article is the preliminary character of its empirical evidence: the pilot results of Section 5 are observations from a single site; controlled quantitative validation is outside the scope of the present article and remains an area for future work.
The author invites collaboration across the OIML community on protocol standardization, cross-site interoperability, and the extension of legal-metrology principles into AI-mediated industrial recording.
References
[1] OIML R 76-1:2006 Non-automatic weighing instruments – Part 1: Metrological and technical requirements. Available from https://www.oiml.org/en/publications/recommendations.
[2] S. Liu et al., “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” in Proc. European Conference on Computer Vision (ECCV), 2024, pp. 38–55.
[3] BIPM/JCGM 100:2008, Evaluation of measurement data – Guide to the expression of uncertainty in measurement (GUM). Bureau International des Poids et Mesures.
[4] Japan Patent Application Publication JP 2023-180853 A, “Data collection system, data collection method, model generation method and foreign-object recovery system,” published 2023-12-21.
[5] OIML R 51-1:2006 Automatic catchweighing instruments – Part 1: Metrological and technical requirements. Available from https://www.oiml.org/en/publications/recommendations.
[6] OIML R 134-1:2006 Automatic instruments for weighing road vehicles in motion and measuring axle loads – Part 1: Metrological and technical requirements. Available from https://www.oiml.org/en/publications/recommendations.
[7] OIML D 31:2023, General requirements for software-controlled measuring instruments. Available from https://www.oiml.org/en/publications/documents.
[8] A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple online and realtime tracking,” in Proc. IEEE International Conference on Image Processing (ICIP), 2016, pp. 3464–3468.
[9] W3C, “PROV-DM: The PROV data model,” W3C Recommendation, 30 April 2013.
[10] S. Haber and W. S. Stornetta, “How to time-stamp a digital document,” Journal of Cryptology, vol. 3, no. 2, pp. 99–111, 1991.
[11] ISO 22095:2020, Chain of custody – General terminology and models. International Organization for Standardization, Geneva, 2020.
[12] E. Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. National Institute of Standards and Technology, Gaithersburg, MD, 2023.
[13] “Update from the Digitalization Task Group,” OIML Bulletin 2026 LXVII(1) 20260107.