In the domain of colexification between vision and cognition, previous studies have identified two distinct typological patterns across languages: one where non-volitional visual verbs such as see in English, which emphasize the result of perception, tend to extend into the cognitive domain (Georgakopoulos et al. 2022); and another where volitional visual verbs such as 看 (kàn: look) in Mandarin, which emphasize the action of perception, are more likely to develop cognitive meanings (Dai and Wu 2024a, b). Based on cross-linguistic evidence, this study proposes a third colexification pathway: volitional visual activity verbs which encode concomitant manner information, can also extend into the cognitive domain, such as angalia and tazama in Swahili, and fiiri in Somali. This paper further argues that: (1) from the perspective of lexiconization patterns, look-type verbs, which denote controllable visual activities, primarily lexify [ACTION] as their core semantic component. In contrast, see-type verbs, associated with uncontrollable visual experiences, encode [ACTION + RESULT], while watch-type verbs, also denoting controllable visual activities, encode [ACTION + MANNER] due to their inclusion of manner-specific information; (2) the three types of visual verbs represent three typologically distinct pathways for vision-cognition colexification: languages such as Mandarin and Korean exemplify the look-type path, English and Spanish illustrate the see-type path, and Swahili and Somali follow the watch-type path; and (3) from a cognitive-logical perspective, Mandarin and Korean reflect an emphasis on perceptual thinking rooted in sensory experience, English and Spanish foreground rational understanding of external reality, while African languages such as Swahili and Somali emphasize evaluative cognition derived from detailed observation and appreciation, arguably aligning more closely with rational cognitive processes.