All authors
feiyang-k avatar

Claude Skills by feiyang-k

github.com/feiyang-k
258 skillsA× 257B× 10 installs89 views
The All Seeing Project Towards Panoptic Visual Recognition And Understanding Of The Open World Arxiv 2308 01907v2A

Use this skill when you want to create a large-scale dataset with region-level recognition and text descriptions for panoptic visual understanding. Avoid it when you do not need region-level understanding or panoptic annotations.

testinggo
0
9
Towards Open World Segmentation Of Parts Arxiv 2305 06914v3A

Use this skill when you need part-level segmentation data for training VLMs to understand object parts and fine-grained visual components. Avoid it when object-level segmentation is sufficient.

researchgo
0
9
Trivialaugment Tuning Free Yet State Of The Art Data Augmentation Arxiv 2103 10158v3A

Use this skill when you want the simplest possible automated augmentation: randomly select one transform and one magnitude per image, no tuning needed. Avoid it when you prefer tuned augmentation policies.

developmentgo
0
9
Unnatural Instructions Tuning Language Models With Almost No Human Labor Arxiv 2212 09689v2A

Use this skill when you want to generate diverse instruction data by prompting an LLM with creative seed constraints for novel instruction types. Avoid it when standard self-instruct provides sufficient diversity.

ai-agentsgo
0
9
Video Chatgpt Towards Detailed Video Understanding Via Large Vision And Language Models Arxiv 2306 05424v2A

Use this skill when you want to generate video instruction data by prompting GPT-3.5 with video descriptions and annotations for video understanding VLM training. Avoid it when you do not work with video or have sufficient video instruction data.

ai-agentsgoapi
0
9
Visual Instruction Tuning Arxiv 2304 08485v2A

Use this skill when you want to generate multimodal instruction-following data by prompting a text-only LLM with image captions and bounding boxes. Avoid it when you have abundant human-annotated instruction data or need the LLM to see actual images during data generation.

ai-agentsgodebugging
0
9
Vokenization Improving Language Understanding With Contextualized Visual Grounded Supervision Arxiv 2010 06775v2A

Use this skill when you want to augment text training data with visual tokens ('vokens') retrieved from an image database to ground language in vision. Avoid it when you do not need visual grounding of text data.

ai-agentsgodatabase
0
9
Wizardlm Empowering Large Language Models To Follow Complex Instructions Arxiv 2304 12244v2A

Use this skill when you want to progressively increase instruction complexity through evolution prompts (deepening, widening, adding constraints). Avoid it when you need simple instructions or cannot afford iterative evolution.

code-qualitygoperformance
0
9