Free human-review protocol

Face swap quality scorecard for photos, videos, and GIFs

Rate one generated output with seven visible criteria, three critical publication gates, and a transparent weighted formula. Then map the review to nine metric families and three published video benchmark protocols without treating incompatible scores as equivalent.

By DeepSwapAI Product TeamProtocol published July 18; benchmark crosswalk reviewed July 28, 2026Human-review protocol

A face swap quality score needs visible criteria and an evidence boundary

This scorecard structures a human review of one specific output. It does not inspect a face biometrically, predict an unseen result, or turn one favorable sample into a claim about an entire model or provider.

7 visible criteriaIdentity, blend, pose, lighting, occlusion, technical integrity, and motion stability.
3 critical gatesPermission, intended identity mapping, and publication disclosure remain separate from the number.
0 uploaded review dataThe page accepts no media file. Ratings, notes, calculation, and exports stay in this browser tab.

Score one output with the same standard

Choose the media type, complete the gates, rate every applicable criterion, and keep a short evidence note for anything another reviewer should be able to find.

Runs locally
Output type
Critical publication gates

A score cannot override an unresolved gate.

Rate the visible output

Use 0 for an unusable failure and 4 only when no material defect is observed in the reviewed sample.

Identity preservationDo the visible eyes, brows, nose, mouth, jaw, age cues, and overall identity remain coherent with the intended reference?24% base weight · Review at normal viewing size, then inspect the face during the hardest visible pose or frame.
Face boundary and blendDo skin texture, face edges, ears, jaw, hairline, and neck transition into the target scene without a pasted-on boundary?16% base weight · Inspect high-contrast edges and the full face perimeter rather than only the center of the face.
Pose and expression coherenceDoes facial geometry remain coherent with the target head angle, gaze, eye state, mouth shape, and expression?14% base weight · Prioritize profile turns, closed eyes, open mouths, smiles, and steep head tilts.
Lighting and color continuityDo exposure, color, skin shading, highlights, and shadows remain consistent with the target scene?14% base weight · Check directional light, colored light, shadow boundaries, and transitions between bright and dark regions.
Occlusion and accessory continuityDo hair, hands, glasses, masks, microphones, foreground objects, and other occlusions stay in the correct visual order?12% base weight · Review every point where an object crosses the face or where the face moves behind another object.
Technical integrityIs the output free from material blur, ringing, block artifacts, tearing, duplicate features, abrupt texture changes, or damaged frames?10% base weight · Inspect at the intended delivery size and avoid treating deliberate motion blur or depth of field as a defect by itself.
Temporal stabilityAcross motion, does identity remain stable without flicker, drift, face loss, sudden geometry changes, or a visible loop seam?10% base weight · Review the full clip at normal speed, then replay the hardest turn, occlusion, cut, and loop boundary frame by frame.

Use the same 0-to-4 meaning for every criterion

The scale describes the strongest defect observed in the sampled material for that criterion. It is not a confidence score, probability, or biometric measurement.

RatingAnchorDefinition
0UnusableA failure or severe defect prevents the intended use.
1Major defectA clearly visible defect dominates normal viewing.
2Noticeable defectA defect remains easy to notice and needs targeted correction.
3Minor defectA small defect is visible on review but does not dominate normal viewing.
4No material defect observedNo material defect was observed in the sampled output for this criterion.

Weights are explicit and motion is only scored when motion exists

Video and GIF use all 100 base-weight points. Photo excludes temporal stability and normalizes the remaining 90 points to a 100-point result.

CriterionBase weightMediaReview question
Identity preservation24%photo, video, gifDo the visible eyes, brows, nose, mouth, jaw, age cues, and overall identity remain coherent with the intended reference?
Face boundary and blend16%photo, video, gifDo skin texture, face edges, ears, jaw, hairline, and neck transition into the target scene without a pasted-on boundary?
Pose and expression coherence14%photo, video, gifDoes facial geometry remain coherent with the target head angle, gaze, eye state, mouth shape, and expression?
Lighting and color continuity14%photo, video, gifDo exposure, color, skin shading, highlights, and shadows remain consistent with the target scene?
Occlusion and accessory continuity12%photo, video, gifDo hair, hands, glasses, masks, microphones, foreground objects, and other occlusions stay in the correct visual order?
Technical integrity10%photo, video, gifIs the output free from material blur, ringing, block artifacts, tearing, duplicate features, abrupt texture changes, or damaged frames?
Temporal stability10%video, gifAcross motion, does identity remain stable without flicker, drift, face loss, sudden geometry changes, or a visible loop seam?

Download the reusable protocol

The JSON contains the complete scale, gates, criteria, formula, decision bands, evidence limits, license, and references. The CSV is a blank seven-row review template, and the BibTeX file provides a stable citation.

Face swap quality metrics answer different questions

A defensible benchmark keeps source-identity transfer, target-attribute preservation, image and video distribution realism, frame-to-frame behavior, output-only assessment, and human review separate. The versioned decision map names what each family compares and the protocol needed before a value can be interpreted.

MeasureComparesUseBoundary
Identity retrieval or embedding similaritySource identity versus the swapped output, using a named face encoder or source gallery.How strongly does the output retain source-identity evidence under the fixed evaluator? Higher similarity or retrieval performance is usually better inside the same protocol.The encoder, crop, gallery, demographics, threshold, and preprocessing affect the value. It is not a biometric verdict for one person.
Frame-wise identity similarity stabilityThe sequence of source-to-output identity similarities across detected output frames under one named face encoder.Does source-identity evidence remain stable as pose, expression, occlusion, and scene conditions change over time? Lower dispersion can indicate greater stability only when the mean identity similarity or retrieval result is reported beside it.A consistently wrong identity can have low variance. Dispersion cannot replace mean similarity or retrieval, and values are not portable across encoders or frame pipelines.
Pose and expression errorTarget pose or expression estimates versus the swapped output under a named estimator.How closely does the output preserve target head pose and expression under the fixed estimator? Lower error is usually better inside the same estimator and parameterization.Values are not portable across different estimators, crops, parameter spaces, landmark conventions, or preprocessing pipelines.
Frechet Inception Distance (FID)A distribution of generated images or frames versus a documented reference distribution.How close are two feature distributions under the fixed feature extractor and sampling protocol? Lower is better only inside the same dataset, extractor, preprocessing, and sample protocol.One face swap output cannot have a defensible FID because a single image is not a distribution. FID does not isolate identity correctness.
Frechet Video Distance (FVD)A distribution of generated video clips versus a documented reference-clip distribution under a named video feature extractor.How close are generated and reference video feature distributions under one fixed clip-sampling protocol? Lower is better only inside the same dataset, feature extractor, clip duration, frame rate, preprocessing, and sample protocol.One output clip is not a defensible distribution-level FVD result. FVD does not isolate source identity, target-attribute preservation, lip synchronization, or a specific temporal defect.
Subject and background consistencyFeature consistency across output frames and against preceding, target, or reference frames.How stable are subject and background features through motion under the fixed video protocol? Higher consistency is usually better under the same implementation.Consistency alone cannot establish correct identity mapping, visual realism, motion accuracy, or publication readiness.
Gaze, eye aperture, lip sync, and motion diagnosticsNamed eye, audio-lip, landmark, or optical-flow behavior between target and output sequences.Which specific dynamic target attributes remain synchronized through the generated video? Direction and unit depend on the named diagnostic; report each separately.These diagnostics are not interchangeable, need a documented video protocol, and do not apply to every media type.
Learned no-reference quality assessmentThe output alone, evaluated by a model trained against ranked human judgments or labeled quality data.How does a trained assessor rank visible output quality when explicit reference media are unavailable? Direction depends on the model output; validate it on a documented held-out distribution.Performance depends on training data and validation scope. An output-only score cannot establish correct source identity transfer by itself.
Structured human scorecardOne visible output against explicit defect anchors, publication gates, and reviewer notes.Which visible defects or unresolved use gates does a reviewer observe in this specific output? Use anchored ordinal ratings; do not reinterpret the result as a probability or biometric score.Subjective evidence about the reviewed sample, not model accuracy, provider ranking, probability, or biometric similarity.

Published benchmark protocols are not interchangeable

The crosswalk below records only the setups described by the primary papers. It does not copy their result tables, claim that every dataset is downloadable, or turn values from different evaluators into one leaderboard.

Protocol and sourceComparison unitReported measuresComparability boundary
CASIA FaceSwappingSource-target video pairs assembled under three factor-isolation protocols. Normal: 4,500 non-overlapping same-ethnicity pairs from normal recordings. Cross-ethnicity: 1,200 pairs spanning both directions between Asian, African, and Caucasian groups. Cross-attribute: 4,300 pairs spanning normal and pose, expression, or illumination variations in both directions.identity retrieval, identity similarity, pose error, expression error, FID, subject consistency, background consistency. No human-review protocol is summarized in this crosswalk.The three slices isolate different factors inside the CASIA setup. Their values depend on that paper's pair construction, models, feature extractors, resolutions, and aggregation.
IDBench-V200 paper-reported real-world source-video and target-image pairs. One 200-pair evaluation set covering small faces, extreme head poses, severe occlusions, complex or dynamic expressions, and cluttered multi-person scenes.ArcFace identity similarity, InsightFace identity similarity, CurricularFace identity similarity, frame-wise identity-similarity variance, pose error, expression error, background consistency, subject consistency, motion smoothness, FVD. The paper reports 19 evaluators using 1-to-5 ratings for identity similarity, attribute preservation, and video quality.The row records the paper-reported protocol only. It does not assert public download availability, reproduce a result, or make its values portable to another benchmark.
CanonSwap VFS benchmark100 source-target pairs sampled from VFHQ; each target uses the first 100 frames and four seconds of corresponding audio. One 100-pair video face swap evaluation set with fixed frame and audio windows.identity similarity, identity retrieval, pose error, expression error, gaze error, eye aspect ratio error, LSE-D, LSE-C, optical-flow temporal consistency, FVD. No human-review protocol is summarized in this crosswalk.The protocol uses paper-specific frame, audio, estimator, and feature-extractor choices. Its values must not be compared directly with IDBench-V or CASIA values.

Download the metric decision map

The JSON preserves all 9 metric families, the 3-protocol crosswalk, prerequisites, interpretation directions, evidence boundaries, and primary sources. The CSV contains the metric-family table; the BibTeX record provides stable attribution.

Use the right evidence for the decision: automated metrics need a fixed dataset, evaluator, preprocessing pipeline, sample count, and repeated outputs. This local scorecard records visible defects in one reviewed result and deliberately does not imitate an automated benchmark.

Make each review reproducible enough for another person to inspect

  1. Choose the output type and identify the sample. Select photo, video, or GIF and record a sample identifier, reviewer, and review date.
  2. Confirm the three critical publication gates. Confirm permission, verify the intended identity mapping, and review whether an AI-content disclosure is needed.
  3. Review every applicable criterion. Rate identity, face boundary, pose and expression, lighting, occlusion, technical integrity, and temporal stability where applicable.
  4. Read the weighted result without overriding the gates. Use the result to prioritize corrections. A numerical score never overrides unresolved permission, mapping, or disclosure gates.
  5. Export the complete evidence record. Download JSON or CSV, retain notes, and repeat the same protocol for each output you intend to compare.
Provider comparisons need more evidence: use identical permitted inputs, repeated independent runs, multiple reviewers, documented viewing conditions, and statistical reporting. One hand-picked output cannot establish average model quality.

Subjective review is useful only when its limits are visible

Boundary

The score is a structured human observation of one reviewed output, not a biometric identity measurement.

Boundary

The rubric does not establish model accuracy, average quality, safety, speed, reliability, or provider superiority.

Boundary

A high score does not replace consent, disclosure, legal, accessibility, or audience-specific review.

Boundary

A cross-provider claim requires the same inputs, repeated runs, independent reviewers, documented viewing conditions, and statistical reporting.

What the number can and cannot establish

Does this upload my media or notes?

No. The page has no media input, and scorecard controls do not send ratings or notes to DeepSwapAI.

Is this a biometric identity score?

No. It is a documented human observation of visible output characteristics, not face-recognition accuracy or embedding similarity.

Can one score rank face swap providers?

No. A defensible comparison needs shared inputs, repeated outputs, multiple reviewers, controlled viewing conditions, and statistical reporting.

Why is temporal stability excluded from photos?

A still image has no frame sequence. Photo mode normalizes the six applicable criteria from 90 base-weight points to 100.

Can I reuse the rubric?

Yes. The published JSON and blank CSV use CC BY 4.0 and include the required attribution statement and a downloadable BibTeX record.

Can FID measure one face swap output?

No. FID compares distributions of generated and reference image sets. One image is not a distribution.

Should identity be compared with the source or target?

Identity retrieval and similarity normally compare the output with the source. Pose and expression preservation normally compare the output with the target.

Can CASIA FaceSwapping, IDBench-V, and CanonSwap values be compared directly?

No. Their pair construction, challenge slices, evaluators, feature extractors, frame or audio windows, preprocessing, and aggregation differ. Reproduce one shared protocol before comparing values.

Generate the smallest representative output first

Choose one permitted target and identity reference, generate one representative result, then return to score it before scaling a batch or long clip.

Open photo face swap