Computer vision is moving from experimental veterinary image analysis toward clinical decision support, yet high reported diagnostic performance does not establish transportability, reliable uncertainty estimation, workflow value, or patient safety. This critical review evaluates veterinary computer-vision evidence from 2017–2026 across acquisition, annotation and ground-truth construction, classification, detection, segmentation, measurement, diagnostic performance, validation, representativeness, distribution shift, explainability, calibration, abstention, workflow integration, human–AI interaction, oversight, safety, and governance. A transparent targeted selection captured 63 records; nine exact duplicates were removed, 54 unique records were screened, six were excluded at title/abstract assessment, 48 underwent detailed eligibility assessment, 11 were excluded, and 37 evidence units were retained. The literature shows widening veterinary applications in radiography, ultrasound, pathology, ophthalmic imaging, and movement analysis, but uneven evidential maturity. Controlled comparisons can yield strong task-level results, whereas independent practice-sourced validation can reveal substantial degradation and clinically important misses. The review therefore proposes that clinical translation be judged across distinct boundaries: input validity, target validity, transportability, uncertainty reliability, workflow fit, meaningful human oversight, and fail-safe response. These boundaries form an analytical architecture, not a validated score. Evidence for external generalization, calibration, human-factor effects, prospective workflow benefit, and patient-safety outcomes remains thinner than evidence for technical performance. Methodological principles transferred from human medical AI also require explicit veterinary validation because species structure, acquisition environments, professional roles, and regulatory contexts differ. Clinical use should therefore remain claim-specific, supervised, and bounded by the populations, devices, reference standards, and settings actually tested. The review does not establish causal effectiveness or universal readiness and emphasizes task-, species-, device-, and setting-specific validation.