Importing Fine-Tuned Models
Prepare a model that you fine-tuned outside OCI Generative AI for import, then host it on a dedicated AI cluster.
You can fine-tune a compatible base model outside the service and import the resulting model from an Object Storage bucket or Hugging Face. Importing requires a complete model artifact that meets the compatibility requirements for its base model.
OCI Generative AI doesn't fine-tune imported models. For training the service's supported pretrained models within the service, see Fine-Tuning the Base Models in Generative AI.
Check Compatibility Before Fine-Tuning
Start with a model listed in Compatible Models for Import. Review its model-family page for requirements and compatible hosting hardware. Keep a copy of the base model's configuration and record the model ID, repository revision, and training and export library versions.
- Retain an architecture compatible with the supported base model. Standard LoRA updates, when merged into the original weight matrices, preserve their dimensions. Other changes, such as adding layers or resizing the vocabulary, require separate compatibility checks.
- Meet the base model's Transformers version requirement. See Understand Configuration Metadata.
- Keep the total parameter count within ±10% of the original base model, as specified in the model-family guidance. This is the total model parameter count, not just the number of trainable LoRA parameters.
- For LoRA fine-tuning, merge one adapter into the base model before deployment. Each imported fine-tuned model can contain the merged weights of only one LoRA adapter.
A model that loads successfully in your training environment still needs to pass import validation. Review the final exported files, including their configuration, weight datatype, and tokenizer.
Understand Configuration Metadata
Hugging Face Transformers is the library used to load and save many Hugging Face
models. The transformers_version field in
config.json records the library version used when serializing
the configuration. Saving with another library version can change this field even
when the model architecture is unchanged.
The model-family pages require the fine-tuned model to use the same Transformers version as the supported base model. Compare the recorded versions before import. If they differ, check the export environment and confirm compatibility with Oracle Support before relying on an export from a different version.
| Field or Setting | What to Check |
|---|---|
architectures and model_type |
Identify the intended model class and family and agree with the exported weights. |
transformers_version |
Compare with the base configuration. This library metadata is separate from the architecture and from the version label you assign to an imported model in the Console. |
torch_dtype or dtype |
Check the precision declared by the exporter and the actual stored tensor datatypes. Editing a JSON value doesn't convert weights. |
| Layer counts, hidden sizes, attention settings, and vocabulary size | Check that the configuration describes the saved tensors and tokenizer. Keep architecture settings unchanged for a LoRA export that only updates the original weight matrices. |
quantization_config, when present |
Verify that it describes the actual export and that the resulting model variant is compatible. Removing quantization metadata doesn't convert a quantized checkpoint. |
Keep the configuration consistent with the model. Don't change a version string solely to pass validation or replace the entire configuration without checking it against the exported weights and tokenizer.
Merge LoRA Weights and Check Export Precision
A LoRA adapter checkpoint contains the adapter's parameters and depends on the
base model. Files such as adapter_model.safetensors and
adapter_config.json alone aren't a complete model for this
import workflow.
- Load the same base model revision used for training and the intended adapter.
- Merge the adapter into the base model. With Hugging Face PEFT, use
merge_and_unload()for a supported LoRA configuration. See Merge the weights. - Check the merged model's datatype before saving. Use the base model's validated precision as the starting point. Training or merging can produce weights in a different datatype.
- Save the complete model and its tokenizer to a new export directory. Retain the original base model and adapter so that you can repeat the export.
- Inspect the saved configuration and tensor datatypes, then reload the export and run representative prompts. Check the saved tensor metadata before any load-time conversion that could hide the original stored datatype.
For example, the Qwen/Qwen3-1.7B base configuration
declares torch_dtype as bfloat16 (BF16).
If your merged export contains float32 (FP32) weights, convert
the actual floating-point tensors to BF16 and save a new export for this base
model. Changing only torch_dtype leaves the stored weights unchanged.
BF16 in this example is specific to the selected Qwen model. Check the base model and its requirements before choosing precision for another model. A quantized training checkpoint also needs an appropriate export procedure; not every PEFT or quantization configuration supports merging.
Prepare the Model Artifact Directory
Package one complete export in Hugging Face format. The directory selected during
import must contain config.json and the files needed to load
the model. Depending on the model, these include:
- Full model weights, such as
model.safetensors, or all weight shards and their index, such asmodel.safetensors.index.json. - Tokenizer files, tokenizer configuration, special-token definitions, and the chat template used for the model, where applicable.
- Generation configuration and any processor files required by the model.
Ensure that every shard referenced by an index is present. Use the tokenizer and chat template that match the exported model. For chat models, review the Hugging Face chat template guidance.
For an Object Storage import, upload the complete export to its own model artifact directory. For a Hugging Face import, make the complete export available in the model repository. Keep different exports in separate directories so that configuration and weight files from different runs aren't mixed.
Import and Host the Fine-Tuned Model
- Review the source prerequisites in Managing Imported Models and the required imported model permissions.
- Import the prepared export from an Object Storage bucket or from Hugging Face.
- Check the imported model's status and resolve any import failure before deploying it.
- Create a hosting dedicated AI cluster for the imported model. Select a compatible hardware unit shape available in the target region. See Hardware Unit Shapes by Region.
- Create an endpoint, then test the model with representative requests before using it in production.
Record the export revision and evaluation results alongside the imported model's OCID. Use a descriptive name and version to distinguish fine-tuned exports.
Troubleshoot an Architecture Validation Error
If model creation fails with the following message, review the complete export:
The imported model architecture is not yet validated for serving.
The message doesn't identify which configuration or artifact difference needs
attention. An unchanged architectures value alone doesn't establish
that the export meets all import requirements.
- Confirm that the original base model is listed as compatible and that the selected artifact directory contains the intended export.
- Compare the original and exported
config.jsonfiles using the checks in Understand Configuration Metadata. Identify differences introduced during saving or merging. - Inspect actual stored tensor datatypes. For a BF16 base model, check whether the export unintentionally contains FP32 weights.
- Confirm that the export includes the merged base-model weights, all referenced shards, and the matching tokenizer. Reload and test the complete export locally.
- After correcting the export, upload it to a separate artifact directory and submit a new import. If validation still fails, contact Oracle Support.
Include the exact error, OPC request ID, failure time and time zone, region, imported model OCID if available, base model ID and revision, configuration differences, observed tensor datatypes, and training, merging, and export library versions. This information helps distinguish an artifact problem from a service validation issue.
Plan Training Data and Evaluation
For external fine-tuning, follow the dataset format and training requirements of your selected model and tools. The dataset requirements for creating a custom model within Generative AI describe that separate workflow.
- Use accurate, representative examples for the task, languages, and response formats that the model must handle. Review labels and remove duplicate or conflicting examples.
- Keep training, validation, and test data separate. Keep related examples together when splitting data to reduce information leakage.
- Compare the base model and fine-tuned model on the same held-out examples. Measure task quality and check for regressions in general responses.
- Repeat representative checks after merging, datatype conversion, and endpoint deployment. A successful import confirms that the import completed; evaluate the model's responses for your intended use case.
If you want to train on OCI, review Fine-Tuning with OCI Data Science AI Quick Actions for its supported models and prerequisites. A model trained in another service must still meet the Generative AI import requirements.