Technical Collaboration

Fine-Tuning MOSS-VL with ms-swift: A Practical Guide from Environment Setup to Doubling Performance Metrics

MOSS-VL has been merged into the ms-swift main branch. This article provides a complete workflow from environment setup to launching LoRA or Full SFT, using publicly available models and datasets without relying on internal file paths.


MOSS-VL has been merged into the ms-swift main branch. This article provides a complete workflow, from environment setup to launching LoRA or Full SFT. All examples use publicly available models and datasets and do not depend on internal file paths.

1. Environment Setup

Linux, Python 3.10–3.12, and a CUDA-enabled PyTorch environment are recommended.

git clone https://github.com/modelscope/ms-swift.git
cd ms-swift
python -m venv --system-site-packages .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
python -m pip install "transformers>=4.57.1,<5" "torchcodec==0.7.0" joblib

torchcodec==0.7.0 corresponds to PyTorch 2.8. If you use another PyTorch version, install the corresponding TorchCodec version.

2. Prepare Multimodal Data

ms-swift uses messages to store conversations and images or videos to reference media files. Videos use the <video> placeholder in user messages:

[
  {
    "messages": [
      {
        "role": "user",
        "content": "<video>\nDescribe the video."
      },
      {
        "role": "assistant",
        "content": "The player makes a jump shot."
      }
    ],
    "videos": [
      "/absolute/path/to/example.mp4"
    ]
  }
]

Absolute paths are recommended for media files. When using relative paths, set ROOT_IMAGE_DIR to the base directory of the media paths. This variable is used for images, videos, and audio. Each <video> placeholder must correspond to one of the videos in videos.

The training commands below use the publicly available lmms-lab/VideoChatGPT:Generic dataset, which can be downloaded and run directly. To train on your own data, simply replace --dataset with a local JSON or JSONL file.

3. LoRA SFT

LoRA is suitable for first validating your data and task. By default, it freezes the vision tower and multimodal alignment module and injects LoRA only into the linear layers on the language side.

source .venv/bin/activate

export VIDEO_MIN_PIXELS=256
export VIDEO_MAX_PIXELS=16384
export FPS=1
export FPS_MAX_FRAMES=256
export CUDA_VISIBLE_DEVICES=0

args=(
  --model OpenMOSS-Team/MOSS-VL-Instruct-0708
  --dataset "lmms-lab/VideoChatGPT:Generic#1000"
  --use_hf true
  --split_dataset_ratio 0.01
  --tuner_type lora
  --target_modules all-linear
  --freeze_vit true
  --freeze_aligner true
  --torch_dtype bfloat16
  --attn_impl eager
  --num_train_epochs 1
  --per_device_train_batch_size 1
  --per_device_eval_batch_size 1
  --gradient_accumulation_steps 16
  --learning_rate 1e-4
  --lora_rank 8
  --lora_alpha 32
  --gradient_checkpointing true
  --vit_gradient_checkpointing false
  --gradient_checkpointing_kwargs '{"use_reentrant": false}'
  --packing false
  --max_length 4096
  --max_pixels 262144
  --eval_steps 50
  --save_steps 50
  --save_total_limit 2
  --logging_steps 5
  --warmup_ratio 0.05
  --dataset_num_proc 4
  --dataloader_num_workers 4
  --report_to none
  --output_dir output/moss_vl_lora
)

swift sft "${args[@]}"

4. Full SFT

Full SFT trains the language model, vision tower, and alignment module. The example below uses DeepSpeed ZeRO-3 across 8 GPUs and explicitly enables non-reentrant Gradient Checkpointing for both the language and vision components.

source .venv/bin/activate

export VIDEO_MIN_PIXELS=256
export VIDEO_MAX_PIXELS=16384
export FPS=1
export FPS_MAX_FRAMES=256
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export NPROC_PER_NODE=8

args=(
  --model OpenMOSS-Team/MOSS-VL-Instruct-0708
  --dataset "lmms-lab/VideoChatGPT:Generic#1000"
  --use_hf true
  --split_dataset_ratio 0.01
  --tuner_type full
  --freeze_llm false
  --freeze_vit false
  --freeze_aligner false
  --torch_dtype bfloat16
  --attn_impl eager
  --num_train_epochs 1
  --per_device_train_batch_size 1
  --per_device_eval_batch_size 1
  --gradient_accumulation_steps 2
  --learning_rate 1e-5
  --gradient_checkpointing true
  --vit_gradient_checkpointing true
  --gradient_checkpointing_kwargs '{"use_reentrant": false}'
  --packing false
  --max_length 4096
  --max_pixels 262144
  --deepspeed zero3
  --eval_steps 50
  --save_steps 50
  --save_total_limit 1
  --save_only_model true
  --logging_steps 5
  --warmup_ratio 0.05
  --dataset_num_proc 4
  --dataloader_num_workers 4
  --report_to none
  --output_dir output/moss_vl_full
)

swift sft "${args[@]}"

save_only_model=true saves only the complete model required for standalone inference, significantly reducing disk usage. However, it does not save the optimizer or scheduler and therefore cannot be used for strict checkpoint resumption. Remove this parameter when resuming training.

5. Training Results Example

We trained on short basketball game videos for one epoch and conducted a fixed evaluation on videos that were not included in training. Only representative metrics are shown below, where the Base model already had some capabilities and continued to improve after Full SFT.

MetricBaseFull SFT 1 epoch
JSON / Schema validity rate98.18%100.00%
Exact match of event type sequence45.45%54.55%
Event type LCS Precision82.35%100.00%
Event type LCS-F173.20%78.57%
Jersey color accuracy34.12%57.65%

The checkpoint produced by this Full SFT training can be loaded directly by Swift, and inference validation completed without runtime errors.

6. Usage Limitations

  • MOSS-VL currently does not support packing=true or padding-free training.

  • Gradient Checkpointing must use use_reentrant=false.

  • Full SFT requires substantial GPU memory, checkpoint storage, and optimizer state storage. Confirm your GPU and disk resources in advance.

  • The number and resolution of video frames directly affect memory usage and speed. It is recommended to first complete a LoRA smoke test with a small sample.

Model: OpenMOSS-Team/MOSS-VL-Instruct-0708

Training framework: modelscope/ms-swift

Adaptation PR: modelscope/ms-swift#9944