Artificial intelligence (AI) is transforming wildlife monitoring through automated analysis of images and acoustic recordings for tasks such as detection, filtering irrelevant events, and species identification. Many current approaches rely on large models trained on extensive datasets and deployed in the cloud, including systems such as MegaDetector or SpeciesNet for camera-trap imagery and BirdNET for avian acoustics. Although accurate under well-represented conditions, their performance often degrades when applied to new locations, species communities, or recording environments, highlighting persistent challenges in model generalization. Consequently, researchers increasingly rely on species- or site-specific models and adaptive strategies such as calibration, domain adaptation, and continual learning. At the same time, there is growing interest in moving computation closer to the sensor. Edge deployments using lightweight models on embedded platforms such as Raspberry Pi, Nvidia Jetson Nano, or AudioMoth enable real-time inference in remote environments, but introduce constraints related to computation, memory, and energy consumption. These trade-offs motivate hybrid edge-cloud architectures in which edge devices perform local filtering while more complex models and analysis remain in the cloud. This mini-review synthesizes advances in AI-based wildlife monitoring across image and audio modalities, focusing on generalization, data imbalance, and deployment on resource-constrained devices. We review emerging solutions including adaptive calibration, continual learning, and self-supervised representation learning, and discuss how multimodal AI and hybrid edge–cloud systems may enable scalable, robust, and context-aware ecological monitoring.