Free human-review protocol
Face swap quality scorecard for photos, videos, and GIFs
Rate one generated output with seven visible criteria, three critical publication gates, and a transparent weighted formula. Then map the review to nine metric families and three published video benchmark protocols without treating incompatible scores as equivalent.
Direct answer
A face swap quality score needs visible criteria and an evidence boundary
This scorecard structures a human review of one specific output. It does not inspect a face biometrically, predict an unseen result, or turn one favorable sample into a claim about an entire model or provider.
Interactive evaluation
Score one output with the same standard
Choose the media type, complete the gates, rate every applicable criterion, and keep a short evidence note for anything another reviewer should be able to find.
Rating anchors
Use the same 0-to-4 meaning for every criterion
The scale describes the strongest defect observed in the sampled material for that criterion. It is not a confidence score, probability, or biometric measurement.
| Rating | Anchor | Definition |
|---|---|---|
| 0 | Unusable | A failure or severe defect prevents the intended use. |
| 1 | Major defect | A clearly visible defect dominates normal viewing. |
| 2 | Noticeable defect | A defect remains easy to notice and needs targeted correction. |
| 3 | Minor defect | A small defect is visible on review but does not dominate normal viewing. |
| 4 | No material defect observed | No material defect was observed in the sampled output for this criterion. |
Rubric dimensions
Weights are explicit and motion is only scored when motion exists
Video and GIF use all 100 base-weight points. Photo excludes temporal stability and normalizes the remaining 90 points to a 100-point result.
| Criterion | Base weight | Media | Review question |
|---|---|---|---|
| Identity preservation | 24% | photo, video, gif | Do the visible eyes, brows, nose, mouth, jaw, age cues, and overall identity remain coherent with the intended reference? |
| Face boundary and blend | 16% | photo, video, gif | Do skin texture, face edges, ears, jaw, hairline, and neck transition into the target scene without a pasted-on boundary? |
| Pose and expression coherence | 14% | photo, video, gif | Does facial geometry remain coherent with the target head angle, gaze, eye state, mouth shape, and expression? |
| Lighting and color continuity | 14% | photo, video, gif | Do exposure, color, skin shading, highlights, and shadows remain consistent with the target scene? |
| Occlusion and accessory continuity | 12% | photo, video, gif | Do hair, hands, glasses, masks, microphones, foreground objects, and other occlusions stay in the correct visual order? |
| Technical integrity | 10% | photo, video, gif | Is the output free from material blur, ringing, block artifacts, tearing, duplicate features, abrupt texture changes, or damaged frames? |
| Temporal stability | 10% | video, gif | Across motion, does identity remain stable without flicker, drift, face loss, sudden geometry changes, or a visible loop seam? |
Download the reusable protocol
The JSON contains the complete scale, gates, criteria, formula, decision bands, evidence limits, license, and references. The CSV is a blank seven-row review template, and the BibTeX file provides a stable citation.
Automated benchmark map
Face swap quality metrics answer different questions
A defensible benchmark keeps source-identity transfer, target-attribute preservation, image and video distribution realism, frame-to-frame behavior, output-only assessment, and human review separate. The versioned decision map names what each family compares and the protocol needed before a value can be interpreted.
| Measure | Compares | Use | Boundary |
|---|---|---|---|
| Identity retrieval or embedding similarity | Source identity versus the swapped output, using a named face encoder or source gallery. | How strongly does the output retain source-identity evidence under the fixed evaluator? Higher similarity or retrieval performance is usually better inside the same protocol. | The encoder, crop, gallery, demographics, threshold, and preprocessing affect the value. It is not a biometric verdict for one person. |
| Frame-wise identity similarity stability | The sequence of source-to-output identity similarities across detected output frames under one named face encoder. | Does source-identity evidence remain stable as pose, expression, occlusion, and scene conditions change over time? Lower dispersion can indicate greater stability only when the mean identity similarity or retrieval result is reported beside it. | A consistently wrong identity can have low variance. Dispersion cannot replace mean similarity or retrieval, and values are not portable across encoders or frame pipelines. |
| Pose and expression error | Target pose or expression estimates versus the swapped output under a named estimator. | How closely does the output preserve target head pose and expression under the fixed estimator? Lower error is usually better inside the same estimator and parameterization. | Values are not portable across different estimators, crops, parameter spaces, landmark conventions, or preprocessing pipelines. |
| Frechet Inception Distance (FID) | A distribution of generated images or frames versus a documented reference distribution. | How close are two feature distributions under the fixed feature extractor and sampling protocol? Lower is better only inside the same dataset, extractor, preprocessing, and sample protocol. | One face swap output cannot have a defensible FID because a single image is not a distribution. FID does not isolate identity correctness. |
| Frechet Video Distance (FVD) | A distribution of generated video clips versus a documented reference-clip distribution under a named video feature extractor. | How close are generated and reference video feature distributions under one fixed clip-sampling protocol? Lower is better only inside the same dataset, feature extractor, clip duration, frame rate, preprocessing, and sample protocol. | One output clip is not a defensible distribution-level FVD result. FVD does not isolate source identity, target-attribute preservation, lip synchronization, or a specific temporal defect. |
| Subject and background consistency | Feature consistency across output frames and against preceding, target, or reference frames. | How stable are subject and background features through motion under the fixed video protocol? Higher consistency is usually better under the same implementation. | Consistency alone cannot establish correct identity mapping, visual realism, motion accuracy, or publication readiness. |
| Gaze, eye aperture, lip sync, and motion diagnostics | Named eye, audio-lip, landmark, or optical-flow behavior between target and output sequences. | Which specific dynamic target attributes remain synchronized through the generated video? Direction and unit depend on the named diagnostic; report each separately. | These diagnostics are not interchangeable, need a documented video protocol, and do not apply to every media type. |
| Learned no-reference quality assessment | The output alone, evaluated by a model trained against ranked human judgments or labeled quality data. | How does a trained assessor rank visible output quality when explicit reference media are unavailable? Direction depends on the model output; validate it on a documented held-out distribution. | Performance depends on training data and validation scope. An output-only score cannot establish correct source identity transfer by itself. |
| Structured human scorecard | One visible output against explicit defect anchors, publication gates, and reviewer notes. | Which visible defects or unresolved use gates does a reviewer observe in this specific output? Use anchored ordinal ratings; do not reinterpret the result as a probability or biometric score. | Subjective evidence about the reviewed sample, not model accuracy, provider ranking, probability, or biometric similarity. |
Published benchmark protocols are not interchangeable
The crosswalk below records only the setups described by the primary papers. It does not copy their result tables, claim that every dataset is downloadable, or turn values from different evaluators into one leaderboard.
| Protocol and source | Comparison unit | Reported measures | Comparability boundary |
|---|---|---|---|
| CASIA FaceSwapping | Source-target video pairs assembled under three factor-isolation protocols. Normal: 4,500 non-overlapping same-ethnicity pairs from normal recordings. Cross-ethnicity: 1,200 pairs spanning both directions between Asian, African, and Caucasian groups. Cross-attribute: 4,300 pairs spanning normal and pose, expression, or illumination variations in both directions. | identity retrieval, identity similarity, pose error, expression error, FID, subject consistency, background consistency. No human-review protocol is summarized in this crosswalk. | The three slices isolate different factors inside the CASIA setup. Their values depend on that paper's pair construction, models, feature extractors, resolutions, and aggregation. |
| IDBench-V | 200 paper-reported real-world source-video and target-image pairs. One 200-pair evaluation set covering small faces, extreme head poses, severe occlusions, complex or dynamic expressions, and cluttered multi-person scenes. | ArcFace identity similarity, InsightFace identity similarity, CurricularFace identity similarity, frame-wise identity-similarity variance, pose error, expression error, background consistency, subject consistency, motion smoothness, FVD. The paper reports 19 evaluators using 1-to-5 ratings for identity similarity, attribute preservation, and video quality. | The row records the paper-reported protocol only. It does not assert public download availability, reproduce a result, or make its values portable to another benchmark. |
| CanonSwap VFS benchmark | 100 source-target pairs sampled from VFHQ; each target uses the first 100 frames and four seconds of corresponding audio. One 100-pair video face swap evaluation set with fixed frame and audio windows. | identity similarity, identity retrieval, pose error, expression error, gaze error, eye aspect ratio error, LSE-D, LSE-C, optical-flow temporal consistency, FVD. No human-review protocol is summarized in this crosswalk. | The protocol uses paper-specific frame, audio, estimator, and feature-extractor choices. Its values must not be compared directly with IDBench-V or CASIA values. |
Download the metric decision map
The JSON preserves all 9 metric families, the 3-protocol crosswalk, prerequisites, interpretation directions, evidence boundaries, and primary sources. The CSV contains the metric-family table; the BibTeX record provides stable attribution.
Five-step evaluation protocol
Make each review reproducible enough for another person to inspect
- Choose the output type and identify the sample. Select photo, video, or GIF and record a sample identifier, reviewer, and review date.
- Confirm the three critical publication gates. Confirm permission, verify the intended identity mapping, and review whether an AI-content disclosure is needed.
- Review every applicable criterion. Rate identity, face boundary, pose and expression, lighting, occlusion, technical integrity, and temporal stability where applicable.
- Read the weighted result without overriding the gates. Use the result to prioritize corrections. A numerical score never overrides unresolved permission, mapping, or disclosure gates.
- Export the complete evidence record. Download JSON or CSV, retain notes, and repeat the same protocol for each output you intend to compare.
Evidence boundary and sources
Subjective review is useful only when its limits are visible
Boundary
The score is a structured human observation of one reviewed output, not a biometric identity measurement.
Boundary
The rubric does not establish model accuracy, average quality, safety, speed, reliability, or provider superiority.
Boundary
A high score does not replace consent, disclosure, legal, accessibility, or audience-specific review.
Boundary
A cross-provider claim requires the same inputs, repeated runs, independent reviewers, documented viewing conditions, and statistical reporting.
- ITU-R BT.500-15: Methodologies for the subjective assessment of the quality of television images - General principles for documented subjective image assessment; the DeepSwapAI rubric is not an ITU laboratory test.
- ITU-T P.910 (10/2023): Subjective video quality assessment methods for multimedia applications - General principles for documented subjective multimedia assessment; the DeepSwapAI rubric is a practical single-review workflow.
- Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark - Defines the CASIA FaceSwapping benchmark and documents identity, target-attribute, distribution, and video consistency protocols.
- DreamID-V: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer - Introduces the 200-pair IDBench-V protocol and reports multi-encoder identity similarity, frame-wise variance, target-attribute, video-consistency, FVD, and human-review measures.
- CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation - Reports video-specific identity, target-attribute, synchronization, temporal consistency, and distribution measures.
- Rank-based No-reference Quality Assessment for Face Swapping - Studies learned output-only face-swap quality assessment trained against ranked judgments.
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium - Introduces Frechet Inception Distance as a distribution-level evaluation measure.
- DeepSwapAI controlled input-readiness research - separate input-only evidence with no generated-output score.
- Versioned public scorecard mirror - the v1.6.2 copy in the open research repository.
- Checksum-verified open research release - manifest, SHA-256 checksums, citation files, JSON, and CSV.
- Software Heritage repository snapshot - independently preserved as swh:
1: .snp: f11738615eea 9f999c8709a0 da73130f0426 6209 - DeepSwapAI claim verification methodology - how product facts, comparisons, and evidence boundaries are reviewed.
Scorecard questions
What the number can and cannot establish
Does this upload my media or notes?
No. The page has no media input, and scorecard controls do not send ratings or notes to DeepSwapAI.
Is this a biometric identity score?
No. It is a documented human observation of visible output characteristics, not face-recognition accuracy or embedding similarity.
Can one score rank face swap providers?
No. A defensible comparison needs shared inputs, repeated outputs, multiple reviewers, controlled viewing conditions, and statistical reporting.
Why is temporal stability excluded from photos?
A still image has no frame sequence. Photo mode normalizes the six applicable criteria from 90 base-weight points to 100.
Can I reuse the rubric?
Yes. The published JSON and blank CSV use CC BY 4.0 and include the required attribution statement and a downloadable BibTeX record.
Can FID measure one face swap output?
No. FID compares distributions of generated and reference image sets. One image is not a distribution.
Should identity be compared with the source or target?
Identity retrieval and similarity normally compare the output with the source. Pose and expression preservation normally compare the output with the target.
Can CASIA FaceSwapping, IDBench-V, and CanonSwap values be compared directly?
No. Their pair construction, challenge slices, evaluators, feature extractors, frame or audio windows, preprocessing, and aggregation differ. Reproduce one shared protocol before comparing values.
Generate the smallest representative output first
Choose one permitted target and identity reference, generate one representative result, then return to score it before scaling a batch or long clip.